the question
Overview
I started this project during my rotational program at BMO U.S. It was not a UX Research team assignment. I wanted to know how AI features affect trust in digital banking and which design principles could support AI without costing people their confidence. I reviewed public surveys, provider publications, and academic literature, then coded the evidence by theme and turned the patterns into testable design hypotheses.
This was a synthesis of published research, not a participant study. Survey figures and provider examples link to their original publications. The design implications are my interpretation.
Proposed next step: a moderated prototype test would compare an AI recommendation with and without an explanation and override, then observe whether people catch a flawed suggestion. Read the validation plan →
one thread, from source to hypothesis
How I Synthesized the Sources
The steps below follow one piece of evidence from the source to the design hypothesis it produced. I treated the rest of the material the same way.
In TD Bank's 2025 U.S. survey, 60% of respondents were comfortable using AI for budgeting, compared with 44% for investing. The same report describes transparency and keeping people involved as foundations for AI in banking. TD Bank survey ↗
The same respondents gave different answers for different tasks, so I coded trust as a property of the task, not a fixed trait of the person. Comfort dropped as the delegated decision became more consequential and harder to undo. The oversight theme in the same report pointed the same way: people want to stay involved where the stakes are highest.
Match the safeguards to the stakes. Low-stakes suggestions, such as budgeting nudges, can appear inline. Consequential actions, such as moving money or investing, should give a plain-language reason, a review step, and a visible way to override.
This is why the principles below form a graduated pattern instead of a confirmation on every action. It is also why the two trust needs differ in how much friction people accept, not in whether they trust AI at all.
Caveat: two tasks from one survey suggest a stakes effect but do not measure one. The validation plan below tests this hypothesis.
what the sources support and what still needs testing
Evidence & Strategy Hypotheses
Comfort depends on the task
Comfort with AI was higher for budgeting (60%) than for investing (44%). Confidence varies with the decision being delegated. TD Bank survey ↗
Human oversight remains important
Trust, transparency, and human involvement appear as foundations for AI in banking. This supports review steps before consequential actions. TD Bank survey ↗
Explanation and override are design hypotheses
I translated the trust concerns into proposed explanations, review steps, and override controls. Their effect on real banking behavior still needs to be tested.
Make the benefit visible, then test it
Showing a concrete outcome, such as money saved, could make an AI feature easier to evaluate. This study did not test messages or measure engagement.
two archetypes, reduced to what they explain
Two Trust Needs
I drafted two behavioral archetypes and journey maps from the secondary research. They are kept here only where they explain a recommendation. Both are illustrative and were not validated with participants.
| Trust need | Where trust is most likely to break | Recommendation it supports |
|---|---|---|
| Speed first Delegates routine tasks such as recurring payments and wants the AI out of the way. | Confirmation steps on low-stakes actions, plus slow or opaque status after money moves | Lightweight handling of low-stakes actions, an override that is always visible, and immediate confirmation of the result |
| Explanation first Wants to understand an action before allowing it, especially when a transfer can't be undone. | A recommendation with no reason given, or no review point before a transfer is submitted | Plain-language reasons, a review checkpoint that flags risky details, and an easy way to opt out |
grouped by how closely each example applies
Competitive Benchmarking
These are provider-reported examples, not outcomes measured in this research. I grouped them by who interacts with the AI, because that determines what carries over to a customer banking experience.
Explore the five provider examples
Customer-facing: most directly relevant
Bank of America: Erica
Bank of America reported more than 3 billion Erica interactions and said over 98% of users find the information they need. Erica report ↗
What transfers: customers will use a conversational assistant for everyday banking. Usage volume shows adoption, though, not trust in consequential actions.
TD Bank: AI Prism
TD said its AI Prism model processes 100 times more data variables to help predict customer needs. At the 2025 launch, personalized marketing was a planned use, not a measured result. TD announcement ↗
What transfers: more prediction means more recommendations, and more recommendations need explaining.
Employee-facing: transfers with limits
Bank of America: AskGPS
Bank of America reported that AskGPS draws on more than 3,200 internal documents to answer employee questions within seconds. AskGPS report ↗
What transfers: answers grounded in identifiable sources could serve as a model for customer-facing explanations. What doesn't: employees are trained, accountable, and able to check answers. Customers bear the financial risk alone and need more support.
Behind the scenes: context, not interface patterns
Visa: fraud detection
Visa said it helped proactively block $40 billion in fraud in fiscal 2023. This is a company-reported network figure, not an outcome measured in this project. Visa announcement ↗
Stripe: Adaptive Acceptance
Stripe describes Adaptive Acceptance as an AI-powered product designed to recover false declines and improve payment acceptance. Stripe technical article ↗
What transfers: these systems act on people's money without a visible interface. The customer meets them only when something is flagged or declined, and that moment is where an explanation and a clear next step matter.
design ideas to validate next
Strategic UX Implications
Safeguards matched to the stakes
Keep low-stakes suggestions lightweight. Add explanation, review, and override as the action becomes more consequential or harder to undo.
Explainable AI with progressive disclosure
Lead with a plain-language reason, offer more detail on tap, and save the technical specifics for people who look for them.
Confirmation checkpoints
Let people review consequential AI-initiated actions before they run, and point out details that look risky.
User override controls
Provide a clear way to undo, adjust, or decline a recommendation, and make sure people can find it.
Visible results
Show what happened after an AI action so people can check the outcome.
how I would test the central hypothesis
Validation Plan
Hypothesis: when an AI recommends a consequential money action, a plain-language reason plus a visible override will help people understand the recommendation and challenge it when it is wrong, without slowing routine use much.
A moderated, within-subjects prototype test. Each participant sees both variants, in counterbalanced order.
8–10 people who bank on mobile, screened to include both high and low comfort with AI. The screener would reuse the budgeting-versus-investing comfort question.
A: an AI recommendation with a single Confirm button. B: the same recommendation with a "Why this?" reason and visible Edit and Decline controls.
Respond to an AI-suggested bill-payment date and an AI-suggested transfer to savings. In one trial the recommendation is deliberately flawed: it schedules a payment after the due date.
Whether participants catch and override the flawed recommendation. Whether they can explain any recommendation in their own words. Time to find the override. Perceived control after each variant. Think-aloud comments at the review step.
Keep explanation and override if Variant B improves catches of the flawed recommendation and the quality of participants' explanations. If it only adds time, simplify the explanation and retest. With this sample size, results are directional, not statistically conclusive.
where the evidence stops
Limitations
- No primary behavioral data: the findings come from published industry reports and competitive analysis.
- Publication bias: provider reports describe their own products and may emphasize favorable outcomes.
- Illustrative archetypes: the two trust needs came from secondary research patterns and were not validated through interviews.
- Fast-changing landscape: the original research and the 2025 examples may not reflect newer AI features or expectations.