At Canada Fintech Forum, RBC and Cohere executives laid bare why AI that works in demos fails at mass customer scale: inference costs, regulatory review backlogs, and legacy systems not built for AI.
Canada's Banks Are Deploying AI for Millions of Customers — But Can't Actually Scale It
By Hector Herrera | September 16, 2026 | Finance
Canadian banks use AI more than almost any other industry on Earth — and yet their executives will tell you in public, with unusual candor, that none of them have fully solved how to scale it. At the Canada Fintech Forum on September 15, RBC and Cohere leadership described three concrete blockers standing between controlled AI demos and mass-customer deployment: inference costs at scale, unresolved security reviews, and legacy infrastructure that was never designed to connect to AI inference systems.
This is worth paying attention to outside Canada. The country's banking sector leads globally in AI adoption — more than 30% of Canadian banking and insurance firms already use AI, among the highest rates of any industry worldwide. The scaling problems they're hitting today are the problems that US and European banks will hit in 2027.
The Three Blockers
1. Inference Costs at Scale
AI models that work beautifully in a controlled demo become expensive at volume. A language model interaction that costs fractions of a cent in a pilot runs differently when it needs to process every customer inquiry, every fraud alert, and every personalized recommendation across a retail banking base of millions.
The math is familiar to anyone who has tried to scale LLM-based features in production: inference cost scales linearly with usage, while customer expectations assume near-zero marginal cost. Banks operate on thin net interest margins. Adding significant per-transaction AI costs without a directly attributable revenue lift is not a winning trade on the income statement.
The solution — batching, caching, distilling smaller task-specific models from large general ones — requires engineering sophistication that most bank technology organizations are building from scratch. The large general-purpose models that make demos impressive are rarely the right architecture for cost-efficient production deployment.
Get this in your inbox.
Daily AI intelligence. Free. No spam.
2. Unresolved Security and Regulatory Reviews
Canadian financial regulators, led by OSFI (the Office of the Superintendent of Financial Institutions — Canada's equivalent of the OCC), require banks to clear AI systems through formal model risk reviews before customer-facing deployment. These frameworks were designed for statistical risk models, not large language models with probabilistic outputs and hallucination tendencies.
The result: banks have AI systems that have passed internal testing and are commercially ready — but are waiting months for regulatory review processes that have not been written yet for their model class. The bottleneck here is not the technology. It is the regulatory review pipeline. Banks cannot ship customer-facing AI faster than OSFI can write and enforce the guidance that governs it.
3. Legacy Core Banking Systems
Canadian banks' core infrastructure was built primarily across the 1980s through 2000s. It was not architected to expose real-time transaction data to external inference systems. Connecting a modern LLM to a bank's core ledger, customer profile data, and transaction history requires integration work measured in engineering years, not months.
This is the constraint that surprises technology executives most. The AI model itself is often the easiest part of the deployment. The hard part is making the bank's data accessible to the model in real time, at production latency, without compromising data governance or triggering regulatory data handling requirements.
The Canadian experience maps onto every regulated financial institution attempting real AI deployment:
- Inference economics — Demo-to-production cost gaps will kill projects that look good on paper.
- Regulatory timing — In the US, OCC and CFPB guidance on LLM-based financial products remains thin; the review queue will become a bottleneck as adoption scales.
- Data architecture — Banks that haven't invested in real-time data infrastructure will hit the legacy integration wall before they hit any model quality limitation.
The banks that will pull ahead in AI deployment are not the ones with the best models. They are the ones that solved inference economics, secured pre-approval for their regulatory review frameworks, and modernized data infrastructure before the deployments started — not after.
What to Watch
Watch whether OSFI publishes updated model risk guidance specifically addressing LLMs before end of 2026 — that would be the single biggest unlock for Canadian banks sitting in the regulatory queue. Also watch whether any major bank announces a cloud-provider inference contract with negotiated unit economics rather than building internal AI infrastructure, which would signal the industry settling on a cost model for production deployment.
Hector Herrera builds production AI systems for financial services clients through Hex AI Systems and writes about finance and AI at NexChron.
Did this help you understand AI better?
Your feedback helps us write more useful content.
Get tomorrow's AI briefing
Join readers who start their day with NexChron. Free, daily, no spam.