Model risk & opacity
Inadequate oversight of model performance and data integrity; opaque outputs that can’t be explained to clients or supervisors.1314 Met with explicit model-risk-management duties.20
Adoption is near-universal and spending is real; measured impact is still early and small.
A compiled, source-verified research digest — every claim cites a downloaded source, every figure is drawn from the data behind it. Not a personal essay.
Generative AI in financial services is past proof of concept and short of transformation. More than half of financial professionals now use the technology,2 and 95% of surveyed wealth and asset managers have scaled it to multiple use cases.10 The measured return is another matter. Only 27% of those same managers report a substantial business impact over one to two years,10 and across industries most firms that report any financial benefit estimate it at low levels.12 The gap matters because the prize on paper is large; McKinsey puts generative AI’s annual value to global banking at $200–340 billion.8 What firms say stands between them and that value is governance and regulatory complexity, the most-cited hurdle in the wealth and asset surveys.10 Robo-advice ran the sector’s first version of this experiment and settled it in a particular direction, gathering roughly $1.2 trillion in assets7 while pure-play economics forced retreats and a pivot to hybrid models that keep a human in the client relationship.18 Regulators have so far policed AI through existing technology-neutral rules, the SEC’s first “AI washing” actions and FINRA’s reminders being the clearest markers,420 while flagging bias, opacity, privacy, and model concentration as the risks finance amplifies,1314 with hallucination raised independently on both consumer and systemic grounds.1621
Start with what is settled. Every major survey of the sector reports majority or near-universal use of generative AI, and the trajectory is steep. NVIDIA’s fifth annual State of AI in Financial Services survey found that “more than half” of financial professionals, 52%, now use generative AI, up from 40% the prior year.2 EY-Parthenon surveyed 100 leading firms, half wealth managers and half asset managers, and found 95% have scaled adoption to multiple use cases.10 Advisors have warmed as fast as their firms. In Advisor360°‘s survey of 300 U.S. advisors, 85% now call generative AI a “help” to their practice, up from 64% a year earlier, while the share calling it a “threat” to their livelihood fell from 21% to 8%.9
The harder finding is the gap between that adoption and demonstrated value. Of the 95% who have scaled, only 27% report a “substantial impact on their firm over the past 1-2 years.”10 EY’s sample is 100 firms and the answers are self-reported, so the figure deserves the caution any survey figure does, but it agrees with the cross-industry baseline. Stanford’s 2025 AI Index, drawing on McKinsey survey data, concludes that “most companies that report financial impacts from using AI within a business function estimate the benefits as being at low levels,” with the most common cost decrease under 10% and the most common revenue increase at or below 5%.12
NVIDIA’s respondents are more bullish. “Nearly 70%” report AI driving a revenue increase of 5% or more, and “more than 60%” report cutting annual costs by 5% or more.2 NVIDIA, though, surveys a self-selected, AI-engaged population of executives, data scientists, and developers, so its impact figures should not be read as industry-representative.
The NVIDIA blog post does not state a respondent count.2 NVIDIA’s own report landing page states the survey covered “over 800 financial services professionals” worldwide; an upstream brief had cited “~600.” This corpus treats “over 800” as NVIDIA’s stated sample as retrieved, and flags the discrepancy rather than asserting either number as definitive.3 The adoption and impact percentages are quoted verbatim from NVIDIA’s reporting; the self-selection caveat above applies to all of them.
Advisors and their firms also read the benefit differently. Three out of four advisors in the Advisor360° survey, 76%, say they have enjoyed immediate benefits from gen-AI-enabled tools,9 a figure that sits in visible tension with EY’s 27%. Both can be true at once. An advisor who saves an hour of administration has an immediate benefit; a firm asking whether the technology changed its business over two years is applying a stiffer test, and about a quarter pass it.
EY’s workforce questions measure the gap from the organisational side, and the two numbers say adoption has run ahead of reorganisation. 97% of firms report minimal headcount changes currently; 68% anticipate substantial workforce transformation over the next five years.10 The technology is in the building while the organisation around it stands unchanged, with the transformation booked to the forecast. Regulators have noticed the same pattern. The OECD’s survey of financial regulators cautions that “pilots and experimentation, or the development of AI-based tools does not equate to deployment.”21
On paper, the prize justifies the budgets. McKinsey estimates that across the global banking sector generative AI “could add between $200 billion and $340 billion in value annually, or 2.8 to 4.7 percent of total industry revenues, largely through increased productivity,” equivalent to 9–15% of operating profits.8 Banking ranks among the highest-impact industries as a share of revenue because so much of its work is knowledge work: customer operations, marketing and sales, software engineering, and risk and legal review.8
The money is moving on that expectation. Among EY’s surveyed managers, 75% are budgeting investments exceeding $11 million, and one-third plan significant resource increases within two years.10 Setting the corpus’s two ledgers side by side puts the bet in plain terms. Three-quarters of these firms are spending at eight-figure scale, while the typical firm that reports any benefit, in the cross-industry data, puts the cost saving under 10% and the revenue lift at 5% or less.12 The sector is paying transformation prices for efficiency returns so far, and the wager is that the returns grow into the spend.
Deloitte’s financial-services survey, covering roughly 540 leaders, shows where the spend concentrates and who says it is paying. It sorts respondents into “pioneers,” the 46% who rate their own gen-AI expertise as high or very high, and “followers,” the remaining 54%. The pioneers lead on every measure the survey asks about; on the sharpest, 47% say ROI exceeded expectations, against 17% of followers, and the chart below carries the full set of pairs.11 One caution on the construction: the pioneer label rests on self-assessed expertise and the ROI figures are self-reported, so the respondents who grade their own expertise highly are also grading their own returns.
What separates the cohorts, on Deloitte’s numbers, is organisational readiness. Followers rate themselves highly prepared on talent at 7% and on risk and governance at 16%, and they most often cite a missing adoption strategy, talent, and governance as the constraint.11 Both cohorts can buy the same models. What the majority of the sector says it lacks is the strategy, people, and controls to deploy them.
The use cases those budgets buy cluster in internal efficiency and client service, and autonomous advice barely features. NVIDIA’s respondents rank trading and portfolio optimisation as the top generative-AI use case by ROI, at 25% of responses, followed by customer experience and engagement at 21%.2 The largest single shift the survey records is customer-facing: chatbot and virtual-assistant use “surged from 25% to 60%,” and more than half of respondents now use generative AI for document processing and report generation.2 Stanford’s function-level data points the same way. The most common cost savings come from service operations, supply-chain management, and software engineering, and the most common revenue gains from marketing and sales and from service operations.12
In wealth itself, appetite runs to the work around the advice. Advisor360° finds demand concentrated in administrative assistance, prospecting, and predictive analytics,9 and its president frames the thesis as “the true promise of AI isn’t in replacing human judgment — it’s in amplifying it.”9 EY sees the next step already underway, with 78% of surveyed wealth and asset managers exploring agentic AI to unlock deeper strategic advantage even as the overwhelming majority keep humans firmly in the loop.10
On the compliance side, deployment is deliberately cautious. FINRA’s 2025 oversight report finds firms “proceeding cautiously,” generally exploring or implementing third-party vendor-supported generative-AI tools “to increase efficiency of internal functions,” which in practice means summarising across multiple sources into one document, conducting analyses across disparate data sets, and helping employees retrieve relevant portions of policies and procedures.19 The same report treats AML and fraud as continuing supervisory focus areas and flags the mirror image, with threat actors exploiting generative AI for “fake content, polymorphic malware, and other malicious tools.”19 The OECD’s jurisdictions see the same thing, reporting generative AI used “as a tool to create sophisticated phishing messages.”21 The technology cuts both ways across fraud detection and fraud commission.
The true promise of AI isn’t in replacing human judgment — it’s in amplifying it.9 — Advisor360° president, AI Connected Wealth Report 2025
The sector has run a version of this bet before, and the result is in. Robo-advice is deterministic portfolio automation, built a decade before large language models, so the technology differs in kind from today’s generative tools. The precedent still carries, because what robo-advice tested was whether automated advice can hold a client relationship at a profit, and that question meets generative deployment in wealth unchanged.
The headline reads as success. The independent Robo Report puts the tracked sector at roughly $1.2 trillion in assets at year-end 2024, up from $1.089 trillion in 2023.7 The composition undercuts the headline. The bulk of the assets sit with incumbents’ hybrid offerings — Vanguard’s combined advice assets of $365.1 billion run more than six times Betterment’s $56.4 billion, the largest of the digital-native pure-plays — and the chart below shows the same shape down the whole table.7 Client growth is slowing even as assets rise. Betterment added clients at roughly 9%, and the other tracked pure-plays grew more slowly still.7
The economics of pure automation have proved hard. JPMorgan shut down its Automated Investing robo in Q2 2024, stating plainly that “the robo-investing business did not take off in the wealth industry as expected. It hasn’t scaled or become profitable for many, including us.”18 The same report catalogues a broader retreat, from Betterment layoffs to BlackRock selling on FutureAdvisor to a Titan SEC fine, and quotes an analyst’s read that robos struggle to “demonstrate sustainable profitability,” with the sector pivoting toward hybrid models that use digital experiences to spark interest and then connect clients with human advisors.18
A behavioural-finance review of 80 peer-reviewed sources explains why pure automation hits a ceiling. It identifies four structural limits: a service-relationship gap, because algorithms cannot build the affective trust central to financial relationships; algorithmic bias, because “algorithms designed by humans … cannot be completely free from human affect, cognition, or opinion”; market-risk persistence, because optimisation within constraints cannot eliminate systematic risk; and financial-literacy stagnation, because passive automation “does not actively engage users in the learning process” and may foster overconfidence.17 The hybrid pivot is, in effect, the industry conceding the first limit. And the first limit is the one that transfers to generative AI whole. However fluent the interface, the affective trust at the centre of a financial relationship is the thing the review finds algorithms cannot build, which is why the robo result reads as a finding about relationship economics and outlives the particular algorithm that produced it.
If governance is what firms say slows them, the regulatory record shows what they are navigating. EY’s managers put regulatory compliance at the top of the hurdle list. 88% of asset managers and 84% of wealth managers name it their greatest hurdle, and the complexity caught 86% of firms by surprise.10 Those are perception measures, and a firm asked to name its greatest hurdle can name the regulator more comfortably than its own talent bench. The record itself, though, supports the complexity claim.
Across jurisdictions the dominant posture is to apply the rules already on the books. The OECD’s survey of 49 jurisdictions finds “the vast majority of respondent jurisdictions have introduced some form of policy that covers AI” in parts of finance, while “only a minority of financial regulators/supervisors … have issued specific” AI guidance; more than a dozen rely on non-binding principles, strategies, and white papers, and the overall approach is “risk-based and technology-neutral.”21 The survey sorts the approaches into three working models: legacy principles-based regimes, as in the UK; AI-specific guidance layered onto existing financial rules, as in Singapore; and cross-sectoral AI regulation integrated into financial supervision, the EU AI Act being the main case.21
Binding legislation is the exception. It exists in a handful of places, the EU AI Act and laws in Brazil, Chile, and Colombia among them, and even the EU AI Act has explicit provisions covering only part of the financial sector.21
FINRA states the U.S. version of the principle directly. Regulatory Notice 24-09 reminds members that existing rules apply to generative AI and large language models, and is explicit that the notice “does not create new legal or regulatory requirements or new interpretations of existing requirements.”20 Rule 3110 on supervision still requires a reasonably designed supervisory system, with policies for AI in compliance review addressing “technology governance, including model risk management, data privacy and integrity, reliability and accuracy of the AI model,” and Rule 2210 applies equally to human- and technology-generated communications.20 A year on, FINRA’s 2025 oversight report adds the supervisory worry in plainer language, finding firms’ use of generative AI “outpacing the controls, documentation and supervisory frameworks needed to manage the technology’s risks.”19
To date the SEC has acted through enforcement. On 18 March 2024 it brought its first “AI washing” cases against two investment advisers. Delphia (USA) Inc. had falsely claimed it used “machine learning” and “artificial intelligence” to analyse client data despite having admitted during a 2021 examination that no such algorithm existed; Global Predictions, Inc. had falsely marketed itself as the “first regulated AI financial advisor” offering “expert AI-driven forecasts.”4 The two settled for civil penalties of $225,000 and $175,000 respectively, $400,000 combined, under the Investment Advisers Act anti-fraud provisions and the Marketing Rule.4 The lesson the cases teach is narrow: the SEC is policing claims about AI, whatever the technology underneath.
The headline rulemaking effort, by contrast, has been abandoned. In August 2023 the SEC proposed rules on “Conflicts of Interest Associated with the Use of Predictive Data Analytics by Broker-Dealers and Investment Advisers” (File No. S7-12-23), which would have required firms to eliminate or neutralise conflicts arising from predictive-data-analytics and AI-driven tools used in investor interactions. On 12 June 2025 the SEC formally withdrew that proposal, one of fourteen rule proposals from the prior administration withdrawn at once, and any future rulemaking on predictive data analytics “must start anew with a new proposal and a fresh opportunity for public comment.”6
The corpus confirms the predictive-data-analytics proposal was formally withdrawn on 12 June 2025.6 It does not establish that AI use in investor interactions is now unregulated: the withdrawal removes a proposed rule, while existing anti-fraud, fiduciary, and Marketing-Rule obligations remain in force, and the source itself “does not explicitly address whether the underlying conflict-of-interest standards already in force remain in effect.”6 The withdrawal ends a specific rulemaking; the existing duties still govern AI advice.
Regulators and standard-setters converge on a short list of risks that leverage, interconnection, and fiduciary duty make sharper in finance than elsewhere. The Financial Stability Board names four vulnerabilities: third-party dependency and concentration risk, correlated market behaviour (“herding”), an expanded cyber attack surface, and model-risk and data-governance failures.14 The IMF reaches a similar list and concludes that generative AI “could aggravate some of these risks and bring about new types of risks as well, including for financial sector stability.”13 The matrix below carries the full mapping of who flags what. The FSB sets those vulnerabilities against acknowledged benefits, from operational efficiency and compliance to customised products and advanced analytics.14 Its recommendation to authorities is to “enhance monitoring of AI developments, assess whether financial policy frameworks are adequate, and enhance their regulatory and supervisory capabilities including by using AI-powered tools.”14
Hallucination gets less attention in the sector’s adoption surveys than in its risk literature, where two sources arrive at it independently. The Roosevelt Institute leads with the consumer side: “Financial institutions’ own uses of Generative AI can hallucinate (that is, produce false or misleading outputs), resulting in harms to their customers, the institutions themselves, and the financial markets in which they operate.”16 The OECD raises the same failure at the system level, warning that “GenAI hallucinations are a concern that could become systemically significant if misleading information” propagates.21 One institution’s fabricated output is a compliance incident; the same failure running through the market’s information supply is a stability problem, and that difference in scale is why a policy institute and a regulators’ survey reached the risk separately.
Lending discrimination is the one harm in this corpus with a direct measurement attached, and the measurement complicates both of the easy stories. The NBER study of U.S. mortgage pricing found that lenders charge Latinx and Black borrowers 7.9 and 3.6 basis points more on purchase and refinance mortgages respectively, roughly $765 million per year in extra interest in aggregate.15 Algorithmic FinTech lenders also discriminate on price, but “40% less than face-to-face lenders,” and the authors find FinTechs do not discriminate in loan approval at all, even though an estimated 0.74–1.3 million minority applications were rejected between 2009 and 2015 due to discrimination overall.15 So the algorithms inherit bias and carry less of it than the humans they replaced. The result supports scrutiny of the data and the model in every deployment, since assuming bias and assuming fairness each gets one half of the measurement wrong.
Opacity collides directly with fiduciary obligation. FINRA’s insistence that AI governance address “reliability and accuracy of the AI model” is, in substance, an explainability requirement, because a firm that cannot explain a recommendation cannot demonstrate it acted in the client’s interest.20 Practitioners feel the adjacent data problem most. Among EY’s surveyed managers, 77% cite concerns around data privacy, accuracy, and external-data use, and the dominant response is defensive, with 87% relying on closed-source models on trusted platforms.10 FINRA’s concrete guardrail runs the same direction, recommending contractual language “that prohibits firm or customer sensitive information from being ingested into a third-party vendor’s open-source Gen AI tool.”19
Concentration is where the risk turns systemic. Then-SEC Chair Gary Gensler warned that institutions will concentrate on a small number of large base models, so that “the whole financial sector, indirectly, will be relying on those central nodes,” and if those nodes “have it wrong, the monoculture goes one way, well, then there’s a risk in society.”5 He framed this as largely beyond the SEC’s reach, since regulators are built around “entities and activities,” and called for diversity of models and data sources to avoid “a pretty fragile system.”5 The Roosevelt Institute sharpens the mechanism for agentic AI. When many agents rely on similar models and data they can react identically, potentially triggering “bank runs and flash crashes,” and dependence on a handful of providers means a single failure could produce “cascading effects throughout the financial system.”16 The same analysis flags a fiduciary trap in the word “agent” itself: despite the name, agents “are not guaranteed to act in users’ interests” and may be designed to “preference the provider.”16
Inadequate oversight of model performance and data integrity; opaque outputs that can’t be explained to clients or supervisors.1314 Met with explicit model-risk-management duties.20
Algorithms inherit bias from designers and historical data.1317 Measured directly in lending — present, but ~40% less than human lenders.15
Firm-level controls address the top row; the bottom-right risk is systemic and needs coordination beyond any one supervisor
Every class of source in this corpus, from vendor survey to consultancy, regulator, and academic review, lands on the same shape. Nearly every firm has adopted, real money is committed, and about a quarter can point to substantial impact. The constraint that evidence supports is organisational and regulatory. Firms name compliance complexity as their greatest hurdle, their preparedness self-assessments say the talent and governance behind it are thin, and the one completed experiment, robo-advice, produced a decade of asset growth without pure-play profitability and resolved itself by putting humans back into the relationship.718
Whether the gap closes from here depends on which of two readings is right, and the corpus gives each one a number. Governance is visibly maturing. In the Advisor360° survey, 82% of advisors say their firms have formal gen-AI policies, up from 47% a year earlier,9 which is what the early stage of a closing gap would look like. The ceiling reading has evidence too. The risks that distinguish finance — opacity against fiduciary duty, measured bias, privacy, concentration — are standing features of the sector,13145 and the robo experiment ran for a decade without moving the structural limit at its centre.1718
The defensible industry-level conclusion is that the technology is past proof-of-concept and short of transformation: real, near-universal, and so far additive at the margin. The economics of advice, planning, and risk have not yet been reshaped.
Whether the adoption-to-impact gap closes as governance matures, or persists because finance’s risk and fiduciary constraints permanently cap what AI can do in client-facing roles, is not resolved by the sources here. The data captures a single cross-sectional moment (2023–2025 surveys); none of the saved sources track the same firms longitudinally from adoption through to measured ROI, which is what resolving the question would require. The agentic-AI wave that 78% of managers are exploring10 is too early in these sources to evaluate on outcomes.