• 4 mins read
  • Published

Banks face real security threats as AI voice agents go live

Peter Warburton Economist and financial markets writer Currency Information

Post by Peter Warburton

Banks face real security threats as AI voice agents go live Currency Information © currencyinformation.org
Banks face real security threats as AI voice agents go live © currencyinformation.org

Banks are rolling out AI voice agents fast, but new tests show standard demos miss big security gaps. Hamming AI's adversarial approach reveals why real-world use needs tougher safeguards.

When a voice AI agent can be fooled into skipping user checks or tripped up by a simple frequency hack, banks face real and immediate risks. At FinovateFall 2026 in New York, Sumanyu Sharma, founder and CEO of Hamming AI, showed just how quickly conversational agents can break down under real-world pressure. Sometimes, the fallout is much worse than a public blunder. This comes as big names like NatWest and Bank of America warn that AI shopping agents can steer users into scams, botch purchases, or leak card data. Reuters has reported on the urgent need for stronger controls in the sector.

Hamming AI is a Silicon Valley fintech that came out of Y Combinator in 2024. The company has tracked over 10,000 conversational agents and run more than 50,000 test calls at once for banks and credit unions. Its platform is built to find weak spots before attackers do. It uses adversarial "red teaming" and constant monitoring to mimic the chaos of real customer calls. This matches the push from regulators like the Bank of England and the European Central Bank, who have called for tougher rules on operational resilience and fraud prevention in recent statements.

Why demos fall short and production is tougher

Many banks judge AI voice agents by slick vendor demos. Sharma says these controlled tests rarely match the messiness of real life. In the wild, agents must deal with all kinds of accents, fast or slow speech, and background noise. Even top models can stumble. Hamming AI's tests push agents far past demo limits, running thousands of fake conversations to hunt for hidden failures. This matters as central banks, including the Federal Reserve, stress the need for strong risk controls as digital payments and systems evolve.

Sharma's team has logged a wide range of flaws across voice, chat, text, and email agents. Modern generative AI models are non-deterministic-they might answer the same prompt in different ways each time. Old-school software testing is no longer enough. Hamming AI uses large-scale adversarial tests, where simulated attackers try to trick agents or break security rules. The Finovate event page notes that Hamming Red Team makes "hundreds of adaptive calls" to both AI voice agents and human call center staff, turning failed calls into regression tests for future updates.

Adversarial attacks and security gaps

At FinovateFall 2026, Hamming AI showed live how a voice agent could be coaxed into skipping required user checks. This exposed the real risk of social engineering and technical tricks. Sharma points out that voice-to-voice models are especially open to acoustic attacks, where attackers inject certain frequencies into the audio to confuse the AI. The fallout isn't just technical. When a voice agent fails, customers often feel more violated than with text-based systems. These worries echo recent warnings from the Reuters banking risk report, which details how AI agents may ask for card details, enter them on websites, or push customers to use payment methods with weaker protections.

To fight these risks, Hamming AI plugs straight into bank operations. It provides real-time monitoring and sets up automated guardrails to block prompt injections, jailbreaks, and accidental data leaks. The system uses a confidence-based queue: clear threats are blocked right away, while unclear cases go to human experts. Every blocked attack is reviewed to make the system stronger. This level of discipline is more important than ever as global authorities like the Bank for International Settlements (BIS) call for better governance and audit trails in AI-powered finance.

Integration and operational impact

Banks often worry that new security tools will be slow or hard to set up. Hamming AI's process is built for speed. Sharma says a basic diagnostic can be running in under ten minutes. Banks just provide a phone number or API endpoint and set the rules. The platform then creates realistic caller personas and launches hundreds of calls at once, producing a detailed report on weak spots and fixes. In some cases, banks have delayed launches after finding critical flaws that standard tests missed. Fast, adaptive testing is now key as central banks watch how AI affects payment system stability and consumer safety. The Federal Reserve's payment systems page gives ongoing updates on risk frameworks.

Hamming AI's compliance setup includes SOC 2 Type II certification and HIPAA-compliant workflows. This puts the company in a strong spot to serve top-100 US banks and credit unions. Its always-on monitoring and fast feedback help banks avoid regulatory fines and keep customer trust by stopping AI agents from making unauthorized promises or leaking sensitive data.

Hamming AI's own numbers show the scale of the challenge. Since launch, it has monitored over 10,000 agents and run more than 50,000 concurrent test calls. These figures show just how much adversarial testing is needed to spot and fix the unique risks of non-deterministic AI in banking. As the ECB and other regulators keep a close eye on AI and financial stability, the demand for advanced red teaming and strong controls is only set to rise.

What red teaming means for AI voice security

Red teaming for AI voice agents means simulating attacks and failures to find weak spots before real attackers do. Unlike old software tests, which expect the same result every time, red teaming for AI must handle the unpredictable nature of generative models. This means building a wide range of test cases-different speech patterns, accents, and background noise-to make sure agents can handle real-world calls. Hamming AI keeps its database of failures up to date and brings in human experts for tricky cases. The goal is to give banks a moving defense against new threats in conversational AI. The focus on controls, governance, and audit at FinovateFall 2026 shows a bigger shift toward active risk management, as top central banks and regulators now demand.

Related Reading