AI Voice Phishing: Success Rate Comparable to Humans—A Single Phone Call Could Cost SMEs “10% of Annual Revenue”
Related Articles
AI Voice Phishing: Success Rate Comparable to Humans—A Single Phone Call Could Cost SMEs “10% of Annual Revenue”
A Single Phone Call, 1,500 Yen. That Could Make Your Company Lose 5 Million Yen
AI has become a fraudster. Moreover, it is as “skilled” as human con artists.
According to a paper published by a research team from the University of Illinois in 2024, the success rate of phishing (vishing) using AI voice agents is nearly equivalent to that of human fraudsters. In a scenario where a GPT-4o-based AI agent tricks individuals into revealing their bank account authentication information, the success rate was approximately 36%. There was no statistically significant difference compared to skilled human fraudsters.
Furthermore, the cost to operate this AI agent just once is a mere $0.75 (about 110 yen).
Hiring a human fraudster would cost several thousand yen per hour. They require training and cannot scale. An AI can make 1,000 calls simultaneously. The total cost would be around 100,000 yen. If even one call succeeds, the fraudster could gain several million yen.
What happens when the cost of an attack drops to one-hundredth? The answer is simple: the number of attacks increases by 100 times. And at the center of this target are small and medium-sized enterprises (SMEs) that lack dedicated security personnel.
The “CEO’s Voice” Can Now Be Copied
Let’s imagine a more specific scenario.
On a Friday evening, the phone rings for the accounting staff. The voice on the other end sounds like the CEO of a business partner. “I need you to make an urgent transfer. It won’t be in time if we wait until Monday. Please process it today.” The tone of voice and speaking style are indistinguishable from the real thing.
In fact, in an incident that occurred in Hong Kong in 2024, a corporate finance officer transferred approximately 3.8 billion yen during a video conference that utilized AI-generated deepfake audio and video. While this was a case involving a large corporation, the structure is the same for SMEs. In fact, SMEs, with their personalized approval processes, may be even more vulnerable.
Current voice synthesis technology can clone a voice with just a 3-second audio sample. Greeting videos on YouTube, seminar presentations, and voice messages posted on social media are all potential samples of a CEO’s voice that are likely already available online.
Estimating the Amount SMEs Could Lose from a Single Phone Call
Let’s consider some specific numbers.
Assuming a Japanese SME (with about 50 employees and annual sales of 500 million yen).
Scenario 1: Transfer Fraud
- Amount of unauthorized transfer: 3 to 5 million yen (monthly procurement payment scale)
- Average time until detection: 2 to 5 business days
- Recovery rate: almost 0% (if transferred to an overseas account)
- Percentage of annual revenue: about 1%
Scenario 2: Chain Damage from Theft of Authentication Information
- Bank account authentication information is leaked → Unauthorized transfer of account balance
- Estimated damage amount: 10 to 30 million yen
- Additionally, secondary damage to business partners and loss of credibility
- Percentage of annual revenue: 2 to 6%
Scenario 3: Induction to Ransomware
- Gaining administrative privileges through voice phishing → System encryption
- Median ransom demand (for SMEs): approximately 5 to 15 million yen
- Lost profits due to business stoppage: 1 million to 2 million yen per day × 5 to 10 days until recovery
- Total estimated damage: 20 to 35 million yen
- Percentage of annual revenue: 4 to 7%
In any of these scenarios, 1% to 7% of annual revenue could vanish in an instant. If multiple scenarios occur simultaneously, it could reach 10%. For a company with 50 employees, a loss of 50 million yen could be fatal. If cash flow becomes tight, bankruptcy could follow.
And the attacker’s cost is just a few hundred to a few thousand yen per incident. This asymmetry is the essence of the problem.
Another Trap—Judgment Errors Due to Overconfidence in AI
It’s not just voice phishing. There is another risk that arises when AI infiltrates the workplace.
A 2024 study from the University of California, Berkeley, found that the decision-making accuracy of subjects who received advice from AI assistants was three times worse compared to when they did not receive any advice. Nevertheless, their confidence in their decisions doubled.
This is not about “AI making mistakes.” It’s about the issue of “when AI makes a mistake, humans may not notice it.”
Consider scenarios that could occur in an SME setting:
- An AI chatbot assesses that “this business partner has a low credit risk” → skips credit review → 5 million yen in bad debt
- AI evaluates that “this candidate has high suitability” → simplifies the interview process → mismatched hiring → 2 million yen loss due to hiring and turnover costs
- Overlooking a clause error in a contract draft generated by AI → litigation risk
The point is not to avoid using AI. The danger lies in treating AI outputs as “correct” without question. AI is a tool for creating drafts, not a party to which final decisions should be entrusted.
Defense Costs Are “Cheaper Than Expected”—Yet Not Implemented
So, how can we protect ourselves?
Let’s organize realistic measures for SMEs and their associated costs.
1. Thorough Callback Authentication (Cost: 0 yen)
“If you receive a call requesting a transfer, hang up and call back the registered number.” This alone can prevent most voice phishing. The cost is zero. All that is needed is strict adherence to the rule. Just put up a piece of paper in the office.
2. Dual Approval Process for Transfers (Cost: 0 yen to a few thousand yen per month)
For transfers above a certain amount, require approval from at least two people. Many online banking settings can accommodate this. The biggest risk is the personalized accounting flow.
3. Phishing Training for Employees (Cost: 100,000 to 300,000 yen per year)
There are training services that use simulated phishing emails and calls. Available from around 10,000 yen per month. Data shows that just conducting this once a quarter can reduce the click rate (the rate of falling for phishing) by an average of 60%.
4. Double-Check System for AI Outputs (Cost: 0 yen)
Do not treat AI outputs as the final decision. Always have a human review them. Just formalizing this rule is enough. Make “AI said it, so it’s correct” a forbidden phrase in the office.
5. Introduction of Voice Authentication and Passphrases (Cost: 0 yen)
Establish a rule within the company that “passwords will be used for transfer instructions given over the phone.” This low-tech solution is extremely effective against AI voice cloning.
In total, the annual cost would be around 100,000 to 300,000 yen. If even one incident costing 5 million yen can be prevented, the ROI would be 15 to 50 times.
The issue is not the cost. It’s the “normalcy bias” that leads to the belief that “we should be fine.”
Defenses That SMEs Can Implement “Because They Are SMEs”
Here, I want to present a reversal perspective.
Large corporations struggle to instill security policies among thousands of employees. It can take months to coordinate between departments. There are three levels of approval for tool implementation.
SMEs are different. If the CEO says in a morning meeting, “From today, we will implement a call-back rule for transfer requests,” it can change that very day. In a company of 50, everyone’s face is visible. There’s a proximity that allows employees to question, “Is this voice really the CEO?”
The speed of decision-making, the small size of the organization, and the visible relationships are structural strengths of SMEs. This strength also applies to security.
While large corporations spend tens of millions of yen to build zero-trust architectures, SMEs can achieve equal or greater defensive effects with a zero-cost rule change.
So, What Should We Do?
Just do three things. You can start tomorrow.
1. Make it a company-wide rule to “call back to confirm transfer instructions given over the phone.” Today.
2. Require dual approval for transfers above a certain amount. By the end of this week.
3. Incorporate simulated phishing training once a quarter. Find a vendor by the end of this month.
In a world where the success rate of AI voice phishing is comparable to that of humans, the habit of “trusting the voice on the phone” has become a vulnerability.
The cost of an attack is 110 yen. The cost of defense can also start from zero.
The problem is not the technology. The most expensive second is the one when you think, “It should be fine for now.”
JA
EN