AI chatbots are quickly becoming a frontline channel for consumer-facing financial services, but new research suggests the technology may be far riskier than many institutions realize. An applied AI researcher who tested 24 leading AI models configured as banking customer-service chatbots found that every single one was exploitable, sometimes with alarming success rates. The findings raise serious implications for banks, fintechs, healthcare providers, and collection operations that rely on chatbots to handle disputes, eligibility questions, and account-related conversations — areas where a single incorrect or misleading response can trigger regulatory exposure or enable fraud.
The analysis, conducted by Milton Leal, lead applied AI researcher at TELUS Digital, underscores a growing reality for regulated industries: when a chatbot gives bad guidance, regulators treat it as a compliance failure, not a technology hiccup.
Leal conducted adversarial testing against 24 AI models from major providers, including OpenAI, Anthropic, and Google, each configured to act as a banking customer-service assistant. All 24 proved exploitable, with attack success rates ranging from roughly 1% to more than 64%.
What went wrong
Across the tests, several troubling patterns emerged:
- Inaccurate or incomplete guidance: Chatbots misquoted eligibility criteria, summarized credit factors that should not be disclosed, or provided misleading information about disputes. In a regulated environment, these responses carry the same legal weight as statements made by trained agents.
- Sensitive information leakage: Some chatbots refused a request but then immediately disclosed sensitive information anyway, a pattern Leal describes as “refusal but engagement.”
- Operational opacity: Many deployments lacked the logging and audit trails regulators expect, making it difficult to reconstruct what happened after a bad interaction.
In some cases, simple prompts extracted proprietary creditworthiness scoring logic or internal documentation meant only for employees. Those same techniques could be reused by fraud rings looking to refine synthetic identity or account takeover schemes.
Regulators are paying attention
Since 2023, the Consumer Financial Protection Bureau has made clear that chatbots must meet the same consumer protection standards as human agents. The Office of the Comptroller of the Currency has echoed that position, emphasizing that AI-driven customer service tools are regulated systems, not experimental pilots.
Why this matters for collections
For organizations deploying chatbots to handle disputes, payment questions, or account information, the takeaway is clear: speed-to-market without strong guardrails creates risk. Leal argues that chatbots should be treated like any other regulated system, with defined ownership, validation, continuous adversarial testing, and clear escalation paths to humans.
As GenAI becomes embedded in high-stakes consumer journeys, the question is no longer whether to use chatbots, but whether organizations can prove that every automated response meets the same compliance standards as a live agent.




