CaraComp
CaraComp
Forensic-Grade AI Face Recognition for:
Get Started7-day refund guarantee**
digital-forensics

Deepfake video call security: closing the $4M verification gap

Your CFO Just Called. It Wasn't Him. $25 Million Is Gone.

A piece of software running on a standard gaming PC, marketed in Chinese-language channels under the warm greeting "HELLO BOSS," can swap a scammer's face for anyone else's, live, in real time, on a Zoom call, and it has already earned its operators an estimated $4 million. Not through some nation-state cyberattack. Not in a Hollywood production house. On consumer hardware, over the same video conferencing tools your finance team used this morning.

TL;DR

Real-time deepfake tools are already running inside live business calls at global scale, platform disclosure labels are a post-incident band-aid, and the only credible defense is moving verification upstream, before trust is granted.

While the tech industry spends its energy debating content disclosure, Instagram is currently testing an optional "AI creator" label, which is a nice idea and nearly useless against fraud, the operational reality has moved somewhere else entirely. 404 Media obtained and tested Haotian AI, the Chinese-developed software now powering impersonation scams across WhatsApp, Zoom, and Microsoft Teams. Their reporters ran it on a live Teams call. It worked. And the leading academic deepfake detector, the one researchers actually trust, failed to catch it.

That's not a gap. That's a structural failure. And it has implications for every fraud investigator, claims handler, compliance officer, and KYC analyst who still operates on the assumption that a video call is corroborating evidence.


How deepfake video calls shifted threat modeling

Deepfake scams moved from post-production to live calls

Deepfake scams used to require editing time. A criminal built the fake video first, then sent it out to trick someone later. That gap gave security teams a window to catch the fraud before real money moved, but that window is now closing fast on live video calls.

Deepfake zoom sessions now bypass old verification habits

A deepfake zoom call looks and sounds exactly like the real employee it is impersonating, which is what makes it so dangerous for finance and security teams. Standard call etiquette, asking someone to turn their head or repeat a phrase, no longer reliably exposes the fraud the way it once did.

Here's the thing people keep missing. Earlier deepfakes, the ones that generated all the congressional hand-wringing a few years back, were post-production artifacts. You made a video, you manipulated it, you distributed it. The manipulation happened before the interaction. That meant two things: the attack was detectable after the fact, and it couldn't pass the "let's hop on a video call to verify" check that fraud teams worldwide had quietly adopted as their last line of defense.

Haotian AI eliminates that limitation entirely. The manipulation happens in real time, inside the call itself, indistinguishable from the live feed. There is no artifact to analyze afterward because there's no post-production step. The fraud is the interaction. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them. This article is part of a series, start with Deepfakes Fool Your Eyes In 30 Seconds The Math Catches Them.

3,000%
spike in deepfake fraud attempts recorded in 2023 alone, as generative AI compressed the attack timeline from weeks of preparation to under 30 seconds
Source: Keepnet Labs

The scale of this isn't theoretical anymore. Deepfake files went from roughly 500,000 in 2023 to a projected 8 million in 2025, according to Keepnet Labs. The FBI's Internet Crime Complaint Center recorded $16.6 billion in losses across AI-assisted fraud in 2024, per Vectra AI. The most cited single incident, a finance employee at engineering firm Arup who transferred $25.6 million after a deepfake video call impersonating the company's CFO, has become the cautionary tale everyone knows but nobody has actually built a defense against.

That Arup employee did exactly what any sensible fraud-aware professional would do: they verified by video. The video lied to them.


Why deepfake fraud labels aren't enough defense

Face swapping tools have outpaced label-based security

Face swapping software like Haotian AI runs faster than the disclosure systems built to flag it, which means the label arrives, if at all, long after the call has ended. Security teams that rely on platform labels to catch face swapping are defending against yesterday's threat with yesterday's tools. Real protection has to sit inside the call itself, not in a badge applied to content after the fact.

Instagram's optional AI creator label is not a bad idea in isolation. For content moderation, for flagging synthetic media in political advertising, for journalism transparency, sure, labels have a role. But fraud doesn't operate on the content discovery timeline. Fraud operates at the moment of trust: the wire authorization, the identity claim, the onboarding session, the insurance submission. By the time any platform label surfaces on content that's already been used to impersonate a CFO on a private call, the money is gone.

"Organizations whose fraud-defense playbooks include 'if it looks suspicious, hop on a video call to verify' are operating on outdated assumptions, the video call is no longer the corroborating signal it was even 12 months ago." Analysis via CyberSignal, on the Haotian AI threat model

The detection problem is arguably even worse. Research published in the Deepfake-Eval-2024 study, cited in a comprehensive breakdown by TrueScreenfound that deepfake detectors collapse below 50% accuracy under real-world conditions. That's below a coin flip. Against a tool like Haotian AI that runs live and adapts frame by frame, static artifact-based detection isn't just inadequate, it's essentially irrelevant. You'd get better results asking the person on screen to hold up today's newspaper.

And yet most enterprises are still not prepared for this. Sumsub's 2026 fraud trends analysis found that the majority of organizations lack formal protocols for handling AI-generated audio and video attacks, even as deepfake fraud scales into subscription-based criminal toolkits. The infrastructure for running these scams, the Haotian AIs of the world, is maturing faster than the institutional response to it.

Why This Matters Right Now

  • โšก The "verify by video call" fallback is deadreal-time impersonation tools mean live video is no longer a reliable trust signal for high-stakes decisions
  • ๐Ÿ“Š Detectors are losing the arms racebelow 50% accuracy in real conditions means automated detection tools cannot be the primary defense layer
  • ๐Ÿ” Fraud teams need new verification architecturemulti-channel callbacks, out-of-band codes, and forensic comparison workflows must replace appearance-based trust
  • ๐Ÿ”ฎ Liveness validation is the emerging standardISO/IEC 30107-3 aligned liveness detection is becoming the floor for remote identity verification in regulated environments

Trusted by Investigators Worldwide
Run Forensic-Grade Comparisons in Seconds
Detailed facial comparison reports. Results in seconds.
Get Started
7-day refund guarantee**

What Verification Actually Needs to Look Like Now

Real-time calls demand layered security, not a single tool

A single security check on a real-time call is only ever one point of failure away from disaster. Layering independent channels around every high-value video call turns a single successful deepfake into just one signal among several that a fraud team can cross-check before releasing funds.

Let's get specific, because this is where most commentary goes vague and unhelpful. "Better verification" means nothing without operational detail. What the Haotian AI situation actually demands is a rethink of the verification stack, not a single upgraded tool, but a layered architecture that assumes any single channel can be compromised. Previously in this series: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Previously in this series: Voice Cloning Identity Fraud Verification Collapse. Previously in this series: Your Voice Is No Longer Proof Youre You And Ghana Just Prove. Previously in this series: Age Estimation Vs Identity Verification Biometric Distinctio. Previously in this series: Your Biometric Age Check Isnt Verifying Identity And Defense. Previously in this series: Mobile Biometrics 2026 Prediction Governance Gap. Previously in this series: Mobile Biometrics Hit The Street In 2026 And The Rules Haven. Previously in this series: Ice Iris Scanner Deployment Mobile Biometrics Field Standard. Previously in this series: Ice To Flood Streets With 1 570 Iris Scanners Heres What It . Previously in this series: Biometric Access Control Decision Stack 2026. Previously in this series: Your Face Unlocks Nothing The 3 Hidden Layers Deciding Who G. Previously in this series: Police Drone Ai Facial Recognition Oversight Gap. Previously in this series: Ai Voice Cloning Silent Call Scams Fraud Investigation. Previously in this series: 3 Seconds Of Audio A 95 Voice Clone Why Investigators Cant T. Previously in this series: Biometric Age Verification Bars How It Works. Previously in this series: Contactless Biometrics Market Growth 2033 Identity Verificat. Previously in this series: Deepfake Takedown Speed Delhi High Court Personality Rights. Previously in this series: Deepfake Detection Geometric Inconsistency. Previously in this series: Deepfake Evidence Courtroom Verification Workflows.

For high-stakes transactions, the new minimum looks something like this: a video call, yes, but cross-referenced simultaneously with a phone callback to a pre-registered number, a confirmation message from the requestor's verified internal account, and for anything above a defined financial threshold, an out-of-band code delivered through a separate pre-authenticated channel. The point isn't paranoia, it's that impersonating someone convincingly across four independent communication channels simultaneously is genuinely hard, even with Haotian AI running on a gaming PC.

For investigators and fraud teams specifically, liveness detection is the piece most organizations are underinvesting in. As Regula Forensics outlines, modern liveness detection, passive, active, or hybrid, is designed specifically to determine whether a submitted biometric sample reflects a genuinely live human presence during a remote session, rather than replayed video, injected media, or a deepfake feed. The goal is to make the verification moment itself strongly resistant to spoofing, rather than trying to detect manipulation after the fact.

This is where facial recognition infrastructure matters, not in the surveillance context that dominates headlines, but in the forensic comparison workflow. When a fraud investigator needs to determine whether the person in a submitted selfie video matches the person on a claimed identity document, that comparison needs to account for the possibility that either the video or the document image has been synthetically generated. Platforms that build facial comparison tools for investigators and fraud analysts are increasingly being asked not just "do these faces match?" but "is there evidence either of these samples was generated rather than captured?" Those are different questions with different technical requirements.

The convergence of liveness validation and facial forensics is where the real defensive work is happening, and it's where the gap between what's possible and what most organizations have deployed is widest.


The Authority Bias Problem Nobody's Talking About

There's a psychological dimension here that deserves more attention than it gets. The reason deepfake impersonation scams are so effective isn't just technical, it's that we are neurologically wired to trust visual and auditory authority signals. A video of your CEO giving instructions activates compliance instincts that a suspicious email never would. Fraudsters running Haotian AI aren't just exploiting a software gap; they're exploiting the same authority bias that makes us follow instructions from someone in a uniform. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone. Up next: Your Cfo Just Called It Wasnt Him 25 Million Is Gone.

That's why the label approach is so fundamentally misaligned with the threat. Labels work on skeptical, information-processing consumers who are already in evaluation mode. Deepfake fraud works on time-pressured professionals who are in compliance mode, responding to an apparent authority figure making an urgent request. By the time an "AI-generated" flag surfaces anywhere in that interaction, the wire has already been initiated.

"Generative AI has compressed what used to take a skilled fraudster weeks of research into a 30-second voice clone and a real-time video filter." Jazz Cybershield, on the compressed deepfake attack timeline

The defense against authority bias isn't skepticism training, it's removing the human judgment call from the highest-risk moments entirely. Mandatory verification protocols that trigger automatically at defined thresholds, regardless of how legitimate the request appears, break the authority bias loop at the system level rather than the individual level. You don't ask employees to be suspicious of their CEO. You build processes where the CEO's video call, however convincing, cannot by itself authorize a wire transfer.

Key Takeaway

Platform disclosure labels address content authenticity after distribution. Deepfake fraud attacks trust at the moment of interaction. Closing that gap requires verification infrastructure, liveness validation, multi-channel callbacks, forensic facial comparison, deployed upstream, before trust is granted, not labels applied after the fact.

The uncomfortable question for every organization running remote operations right now is this: your verification protocols were built on the assumption that a video of a person is evidence of that person. Haotian AI, available now, earning millions, undetectable by the best academic tools in real conditions, it has already made that assumption false. So when a voice note, a selfie video, and an ID image can all be synthesized convincingly in real time, what combination of independent signals would you actually be willing to call court-admissible proof of identity?

If you don't have a specific, documented answer to that question, you don't have a fraud defense. You have a disclosure policy, and those are very different things.

Security teams evaluating deepfake video call defenses should treat this as a budget and process question, not just a technology purchase. Real protection against a deepfake video call comes from combining callback verification, out-of-band codes, and liveness checks into a single mandatory workflow for anything above a set dollar threshold. Facial biometrics systems, when paired with liveness detection, can help flag whether a submitted video call session shows signs of face swapping rather than a genuine live human.

Detecting deepfakes reliably in a live video call setting still lags behind the tools criminals use to create them, which is exactly why detect deepfakes efforts alone should never be an organization's only line of defense. Security budgets built only around detection software leave the actual moment of financial authorization exposed, since a real-time call can pass a detector one moment and still be fraudulent the next. Layering security controls around the call itself, rather than trusting the call alone, closes that gap.

Digital transformation has moved identity verification onto video calls, chat apps, and mobile onboarding flows, and every one of those digital channels now carries the same deepfake exposure. Security leaders should map every point in their process where a video call or voice call alone can trigger a payment, an account change, or a data release, then add a second, independent channel at each of those points. This kind of security review does not require new hardware; it requires rewriting a checklist and enforcing it without exception, even for calls that appear to come from senior leadership.

Protection also means training staff to recognize that a normal-sounding, normal-looking video call is no longer proof of anything by itself. Combined with callback verification and out-of-band codes, this layered security approach gives fraud teams a fighting chance against tools like Haotian AI, which can defeat both human judgment and automated detection in the same live call.

A security team that wants to stop a deepfake video call before it causes a loss needs to treat every high-value call as unverified by default, not just the ones that feel suspicious. This is a security posture change, not a one-time fix: the security policy has to apply the same way on a calm Tuesday as it does during a rushed, high-pressure request. Building that habit into daily security practice is what actually protects the wire transfer, not the presence of a detection tool on someone's laptop.

Video calls remain useful for plenty of ordinary business, and most video calls will never involve a deepfake at all. But because a deepfake video call is now cheap to run and hard to catch in the moment, every video call tied to money movement, account access, or sensitive data needs the same layered checks described above, regardless of who appears to be on screen. Treating video calls as one input among several, rather than as final proof, is the practical shift security teams need to make this year.

A teams call carries the exact same exposure as a Zoom call, since Haotian AI was demonstrated live on a Microsoft Teams call by 404 Media's own reporters. Any organization that leans on Teams for internal approvals should apply the same callback-and-code requirement to a teams call that it applies to external video call requests, because the software behind the fraud does not care which platform carries the video.

Real-time calls of every kind, voice, video, or screen-share, now sit inside the same threat model, which is why real-time calls need real-time safeguards rather than after-the-call review. A callback placed during the real-time calls session, using a number pulled from an internal directory rather than one supplied by the caller, remains one of the simplest checks available and it works regardless of how convincing the face or voice on the call sounds.

Deepfake detection tools still have a role even though they are not sufficient on their own; running deepfake detection alongside callback verification and liveness checks gives a fraud team more than one chance to catch a fake before money moves. Vendors building deepfake detection into video platforms should be judged on how their tools perform in live, adversarial conditions, not just in clean lab tests, since the Deepfake-Eval-2024 findings show how far detection accuracy can fall in the real world.

Deepfake scams like the Haotian AI operation succeed because they target the moment of decision, not the technology stack, which is exactly why deepfake scams keep working even as detection software improves. Training staff to recognize the pattern of deepfake scams, urgency, authority, and a request to bypass normal approval steps, gives people a second reason to pause, even when the face and voice on the call look completely genuine.

A deepfake zoom call and a deepfake video call are the same threat wearing different platform branding, and any policy written narrowly around one video conferencing tool will miss the other. Fraud teams should write verification requirements at the level of "any video call," not at the level of a specific app, so a deepfake zoom session and a call on another platform both trigger the identical callback and code requirement.

Face swapping is the specific technique behind Haotian AI, and understanding face swapping as a live, frame-by-frame process rather than a one-time edit explains why standard etiquette checks like asking someone to turn their head no longer reliably expose it. Security awareness materials that still describe face swapping as a post-production trick are teaching a version of the threat that no longer matches how it operates on a live call.

Deepfake video call incidents like the Arup case show what happens when verification stops at the video call itself, which is why every deepfake video call policy should specify the exact second channel required before funds move. Organizations that have not yet written a deepfake video call policy should treat this as the immediate gap to close, given how cheaply the underlying software can be run.

Facial biometrics used for onboarding and identity checks need to be paired with liveness detection to have any real value against Haotian AI, since facial biometrics alone cannot distinguish a live captured face from a well-executed swap. Vendors offering facial biometrics for remote verification should be asked directly how their liveness layer performs against real-time face-swap software, not just against static photo spoofing.

Detect deepfakes tools built for enterprise use are only as good as the conditions they are tested in, and organizations that plan to detect deepfakes reliably need to budget for continuous retesting as the underlying generation software improves. A detect deepfakes program that is not updated regularly will fall behind the same way academic detectors already have against Haotian AI.

Call verification protocols written for phone fraud years ago need updating for the video era, since a call today can carry a synthetic face as easily as a synthetic voice. Any call that requests a payment, a password reset, or an account change should be treated as unverified until a second channel confirms it, whether that call arrives as audio, video, or a mix of both.

Frequently asked questions

What is a deepfake video call?

A deepfake video call is a live video conversation, such as one on Zoom, WhatsApp, or Microsoft Teams, where software swaps a scammer's face in real time for someone else's face, making the caller appear to be a trusted employee or executive. Unlike older deepfakes made through editing, the manipulation happens live during the interaction itself, leaving no separate file to analyze afterward.

How much money have deepfake video call scams caused?

One software tool called Haotian AI, marketed in Chinese-language channels, has reportedly earned its operators an estimated $4 million through impersonation scams run over Zoom, WhatsApp, and Microsoft Teams. Separately, a finance employee at engineering firm Arup transferred $25.6 million after being deceived by a deepfake video call impersonating the company's CFO.

Can detection software or labels stop a deepfake video call scam?

Detection and labeling fall short because the leading academic deepfake detector failed to catch Haotian AI running on a live Teams call, and disclosure labels like Instagram's optional AI creator tag arrive, if at all, after the call has already ended. Since fraud happens at the moment of trust, protection needs to be built into the verification process itself rather than applied afterward.

Ready for forensic-grade facial comparison?

Full forensic reports with detailed similarity scoring. Results in seconds.

Run My First Search