CaraComp
CaraComp
Forensic-Grade AI Face Recognition for:
Get Started7-day refund guarantee**
digital-forensicsBy Cara Candelario

Deepfake Fraud Detection: Identity Verification Beats Video Review

"I Saw It on Video" Is Now the Most Dangerous Phrase in Any Investigation
A video call composite illustrates how a deepfake detection system analyzes facial and behavioral signals to expose synthetic media.

In February 2024, a finance employee at Arup, one of the world's most respected engineering firms, sat in on a video conference call with the company's CFO and several senior colleagues. The conversation felt normal. The faces were familiar. The voices were right. At the end of the call, he authorized a wire transfer of $25 million.

None of those executives were real. Every face on that call was a deepfake, a real-time synthetic composite built from publicly available footage. The entire "meeting" was a coordinated fraud. And the thing that made it work wasn't sophisticated hacking or elaborate social engineering. It was the simple, ancient human instinct to trust what we see.

TL;DR

Modern deepfakes are good enough to fool trained humans in real time, which means video and audio are now evidence of appearance, not identity, and every investigator needs structured facial comparison to tell the difference.

Arup Deepfake Proved the Myth: Video Isn't Proof

Here's why this misconception is so forgivable: it was correct for most of human history. For decades, video was expensive to produce, nearly impossible to fake convincingly, and treated by courts as strong corroborating evidence. Investigators built entire methodologies around visual observation. If you saw someone's face on a recording, especially a live feed, that was about as close to certainty as the job allowed.

That era ended quietly, and most people missed it.

The tools required to clone a convincing likeness are now accessible, fast, and cheap. Security researchers at McAfee confirmed that as little as three seconds of audio is enough to produce a voice clone that most listeners cannot distinguish from the real person. Three seconds. That's shorter than a sneeze. And voice is actually the harder problem, faces are easier, because the training datasets are larger and the visual outputs are easier to evaluate during generation.

1,740%
increase in deepfake incidents reported across North America This article is part of a series, start with Age Assurance Becomes The New Kyc And Your Next Ca.

That number isn't a rounding error. A 1,740% increase means deepfake fraud has gone from a theoretical threat to a primary attack vector inside the span of a few years. And according to the Deloitte Center for Financial Services, AI-generated fraud in the United States alone is projected to reach $40 billion by 2027. This isn't a niche problem for cryptocurrency exchanges or offshore transactions. It's showing up in boardrooms, legal proceedings, and insurance claims.


Why Your Brain Can't Catch This, Even When It's Trying

The uncomfortable truth is that the human visual system was not built for this problem. We evolved to recognize faces in real-world conditions: consistent lighting, three-dimensional depth, micro-expressions tied to genuine emotion and muscle movement. A deepfake exploits the fact that most of those signals, when rendered convincingly in 2D video, become indistinguishable from real footage to an untrained eye.

And "untrained" doesn't mean inexperienced. It means human.

Even investigators who've spent careers reading body language and spotting inconsistencies can't reliably identify high-quality deepfakes in real time. The brain isn't running a pixel-level forensic analysis during a video call, it's pattern-matching against memory, filling in gaps with expectation, and generally trying to keep up with the conversation. Deepfake generators are specifically optimized to stay within the envelope of what looks "normal enough" to pass that fast, intuitive check.

"Employees are often untrained to identify deepfakes, particularly as they become increasingly sophisticated, and traditional verification methods, such as matching a name to a face, fail when the face and voice are synthetic." CompassMSP: AI-Generated Deepfakes Are Here to Stay

Here's an analogy that makes this click: think about how banks used to verify identity by asking customers to answer a security question, your mother's maiden name, your first pet. That worked fine when the threat was a stranger guessing. The moment criminals could look up those answers online, the system collapsed. The verification method hadn't changed. The threat had. We're in exactly that moment right now, with video as the "security question."


Trusted by Investigators Worldwide
Run Forensic-Grade Comparisons in Seconds
Detailed facial comparison reports. Results in seconds.
Get Started
7-day refund guarantee**

Why Deepfake Video Calls Beat Every Detection System

So if human observation fails, can AI detection save us? Partially, but not in the way most people hope.

A Purdue University evaluation of 24 commercial, government, and academic deepfake detection systems found meaningful variation in accuracy across tools, and critically, no single-layer system performed reliably across all attack types. The best-performing systems didn't just analyze pixels. They operated across multiple detection layers simultaneously: behavioral signals (does the blink rate match what's expected?), integrity signals (are compression artifacts consistent with genuine capture?), and perceptual signals (do fine details like skin texture and hair hold up under frame-by-frame analysis?). Previously in this series: Biometric Age Checks Deepfake Fraud Investigators .

That level of checking matters. A deepfake that passes a pixel-level check can fail a behavioral check. One that survives behavioral analysis might collapse under compression artifact inspection. No single test is sufficient, and this is exactly why investigators who rely on a single "does it look real?" judgment are working with an incomplete toolkit.

The academic research reinforces this. A peer-reviewed forensic survey published in PubMed Central found that pixel-level analysis alone, the approach that most resembles "looking carefully at the video", is among the least reliable methods for detecting modern synthetic media. The limitations compound when you factor in real-world variables: compressed video from a messaging app, inconsistent lighting, and low-resolution source material all degrade the signals that detection systems rely on.

What You Just Learned

  • 🧠 Three seconds of audiothat's all a voice clone requires, per McAfee security research. Less time than it takes to say your own name twice.
  • 🔬 Multi-layer detection is mandatorybehavioral, integrity, and perceptual checks must run simultaneously; any single-layer approach can be beaten.
  • 💡 Video is now a claimnot a confirmation. It proves someone appeared on screen, not that the person on screen was who they claimed to be.
  • 🎯 The $25M Arup attacksucceeded not because the victim was careless, but because the deepfakes were credible enough to pass live, real-time human scrutiny.

The New Standard: Treating Video as a Hypothesis

This is where the investigator's job fundamentally shifts, and where the real aha moment lives.

A forensic document examiner doesn't look at a signature and ask "does that look right?" They compare it against multiple exemplars of known provenance, check for pen pressure consistency, examine paper fiber compression, and verify ink composition. The question isn't "does it look like a signature?", it's "can I establish, through independent evidence, that this specific person made this specific mark at this specific time?" Up next: Human Deepfake Detection Accuracy 55 Percent Struc.

That's now the standard for video. A recording of someone's face is a starting hypothesis, not a closed finding. The question shifts from "does this look like them?" to "can I verify this is them through structured comparison against reference imagery of known provenance, corroborated by out-of-band evidence?"

Structured facial comparison, the kind that uses Euclidean distance analysis across verified facial landmarks rather than visual impression, becomes essential precisely because it's immune to the cognitive shortcuts that deepfakes exploit. As the team at CaraComp has documented, even a 99% accuracy claim means something very specific and contextualthe threshold settings, reference quality, and dataset composition all determine whether a comparison is actually meaningful. Facial comparison done right isn't eyeballing; it's measurement.

And measurement can be cross-checked. A video call cannot.

The multi-factor framework that emerges from this looks like: structured facial comparison against reference imagery with verified chain of custody + behavioral corroboration (does the communication pattern match known behavior?) + out-of-band confirmation through a separate, previously established channel. Three independent verification streams, none of which a deepfake can simultaneously fake across all three without leaving forensic traces in at least one.

Key Takeaway

Video and audio now prove that someone appeared to be present, not that they actually were. Every visual in an investigation is a claim that requires corroboration through structured facial comparison and independent verification, not just the judgment that it "looked real."

The skills that experienced investigators built over decades, reading behavior, spotting inconsistencies, building contextual understanding, aren't obsolete. They're just incomplete on their own now. The game didn't get easier. It added a new opponent that your eyes can't see.

So here's the question worth sitting with: When you get a key photo or video in a case, what's the first thing you do to convince yourself it's really the person you think it is? If the answer is "I look at it carefully", that answer just became the most expensive three seconds in your investigation.

Deepfake Fraud Detection Starts With Identity Verification, Not Video Review

Deepfake fraud detection is the practice of confirming that the person shown or heard in digital media is who they claim to be, using evidence beyond what the video or audio itself appears to show. This matters because fraud built on synthetic media does not rely on hacking systems, it relies on hacking trust in a face. Organizations that build deepfake fraud detection into their identity verification workflow treat every video call, voicemail, and recorded statement as a claim to be checked, not a fact to be accepted. That shift in posture is the single biggest predictor of whether an organization catches a synthetic media attack before or after the money moves.

Face Swapping Techniques Explain Why Detection Requires Multiple Signals

Face swapping is one of the core techniques behind modern deepfake fraud: a generative model maps one person's facial movements onto another person's likeness in real time, frame by frame. Because face swapping tools are trained on large datasets of real facial expressions, the resulting video can preserve blinking, mouth movement, and head tilt well enough to defeat casual observation. This is exactly why detection systems built for deepfake fraud protection layer behavioral, integrity, and perceptual checks rather than trusting any single visual cue.

Liveness Detection Adds a Layer That Recorded Video Cannot Fake

Liveness detection asks a different question than facial comparison. Instead of asking "does this face match a known reference," liveness detection asks "is there a real, physically present human generating this signal right now." Techniques include prompting for spontaneous movement, checking for depth and reflection consistent with a live camera, and analyzing micro-timing between audio and video that pre-rendered synthetic media struggles to reproduce convincingly. Liveness detection does not replace structured facial comparison, it complements it, closing a gap that face-matching alone leaves open.

Synthetic Media Has Moved From Novelty to Primary Fraud Vector

Synthetic media covers any audio, image, or video generated or substantially altered by AI rather than captured directly from reality. The Arup case shows what happens when synthetic media crosses from entertainment and marketing into financial fraud: a fabricated CFO, fabricated colleagues, and a fabricated meeting were enough to move $25 million in a single wire transfer. As generative tools keep improving, the line between authentic and synthetic media will keep getting harder for the unaided eye to find, which is exactly why structured detection matters more than instinct.

Building an Organizational Defense Against Deepfake Fraud

An effective defense against deepfake fraud combines technology, process, and training rather than relying on any single safeguard. Organizations need identity verification steps that do not depend solely on visual or audio confirmation, out-of-band authentication for high-value transactions, and staff training that explains why a familiar face on a screen is no longer sufficient proof of identity. Risk teams that treat deepfake fraud as a payments-process risk, not just an IT security risk, are better positioned to catch attacks before funds leave the organization. Access controls that require a second, independently verified confirmation for large transfers close the exact gap that the Arup attackers exploited.

Digital transformation has made organizations more efficient, but it has also widened the attack surface that deepfake fraud can exploit, every video call, every recorded authorization, every digital signature is now a potential target. Fraud teams investigating a suspected deepfake fraud incident should treat the recording itself as evidence to be tested, not a record to be trusted, and should document exactly which detection layers were applied and what each one found. Detection of deepfake fraud works best as a standing process built into approval workflows, not a one-time check performed only when something already feels wrong. Real-time fraud detection systems that flag unusual transaction patterns alongside identity anomalies give risk teams a second, independent signal beyond what any single video call can show. As deepfake technology keeps advancing, the organizations that treat digital identity as something to be continuously verified, rather than assumed from a familiar face and voice, will be the ones that catch the next Arup-style attempt before the wire transfer clears.

Fraud investigators who specialize in deepfake fraud detection often describe the discipline in terms borrowed from other forensic sciences: every piece of digital evidence carries a risk score, not a verdict. A video that looks convincing on first viewing might still carry a high fraud risk once it is checked against reference imagery and out-of-band confirmation. Treating risk as something to be measured, rather than felt, is what separates a structured deepfake fraud detection program from one that still depends on a reviewer's gut reaction.

Identity verification protocols exist precisely because a face on a screen is not, by itself, sufficient proof of who someone is. Deepfakes exploit weak identity verification protocols by targeting the exact point where an organization stops checking and starts trusting, the moment a familiar voice or face causes a reviewer to skip the extra step. Strengthening identity verification means adding structured checks at that exact moment, not removing human judgment but backing it up with evidence the eye cannot gather on its own.

Deepfake identity attacks differ from older forms of impersonation because they do not require the attacker to obtain a stolen credential or a hacked account. A deepfake identity can be built from public video and audio alone, which means the attack surface includes every executive, employee, or public figure who has ever appeared on camera. This is why identity verification programs built for deepfake fraud detection assume that any face or voice, no matter how senior or familiar, could be synthetic until confirmed otherwise.

Deepfake attacks on financial workflows tend to follow a recognizable pattern: a request for urgency, a familiar authority figure, and a channel, video or voice, that historically inspired unquestioned trust. Detection programs that map this pattern in advance can flag deepfake attacks before the transaction completes, because the request itself, not just the video quality, contains warning signs. Recognizing the pattern of deepfake attacks is often faster and cheaper than trying to forensically analyze every frame of video in real time.

Deepfake video is now cheap enough to produce that attackers no longer need a specific reason to target a given organization, the tools work at scale. A single piece of deepfake video, run through a generative model, can impersonate a CFO, a vendor contact, or a family member with roughly the same technical effort. This scalability is exactly why deepfake fraud detection has moved from a specialized forensic skill to a standard control that every organization handling wire transfers needs in place.

Deepfake technology continues to improve at a pace that outstrips most organizations' training cycles, which means a defense built around recognizing today's deepfake technology will likely be outdated within a year or two. The more durable approach treats deepfake technology as a moving target and builds detection around verification principles, structured comparison, out-of-band confirmation, liveness checks, that do not depend on knowing exactly how the next generation of tools will work. Investing in principles over pattern-recognition is what keeps a deepfake fraud detection program useful as deepfake technology evolves.

Deepfake scams targeting everyday consumers follow a smaller-scale version of the same playbook used against Arup: a familiar voice, an urgent request, and a channel the victim has learned to trust. Romance scams, fake grandchild emergency calls, and fraudulent job offers increasingly use deepfake scams built from a few seconds of publicly posted audio or video. The same detection principles that protect a corporate wire transfer, structured verification and out-of-band confirmation, apply just as well to an individual deciding whether to send money after an unexpected, urgent call.

Generative adversarial networks (GANs) are one of the core technologies behind realistic deepfakes: two neural networks compete against each other, with one generating synthetic images and the other trying to detect the fakes, until the generator produces output convincing enough to fool its own detector. Understanding generative adversarial networks helps explain why deepfake quality keeps improving so quickly, the technology is, by design, built to defeat detection. This is also why detection systems built for deepfake fraud detection use multiple types of checks together rather than relying on any single test the way an internal GAN detector does.

Detection of fraud built on synthetic media works best when it is layered rather than singular. No individual signal, facial comparison, liveness detection, behavioral analysis, or out-of-band confirmation, catches every attack on its own, but combined they close most of the gaps that a determined attacker could otherwise find. Organizations that invest in detection as an ongoing, layered process rather than a one-time checklist are the ones most likely to catch the next synthetic media attack before money leaves the building.

Frequently asked questions

What is a deepfake detection system and why is it needed?

A deepfake detection system analyzes media across multiple layers, behavioral signals like blink rate, integrity signals like compression artifacts, and perceptual signals like skin texture, to determine whether video or audio is synthetic. It is needed because modern deepfakes, like the one used in the $25 million Arup fraud, can fool trained humans in real time, meaning video alone can no longer serve as proof of identity.

Can a deepfake detection system reliably catch every fake video?

No single-layer deepfake detection system performs reliably across all attack types. A Purdue University evaluation of 24 commercial, government, and academic tools found meaningful variation in accuracy, and pixel-level analysis alone is among the least reliable methods. The best results come from combining behavioral, integrity, and perceptual checks simultaneously rather than relying on one test.

Why can't humans detect deepfakes without a detection system?

Human vision evolved to judge real-world faces using lighting, depth, and genuine micro-expressions, none of which apply cleanly to rendered 2D video. Even trained investigators can't reliably spot high-quality deepfakes in real time because the brain pattern-matches quickly instead of running forensic analysis, and deepfake generators are built specifically to pass that fast, intuitive check.

Ready for forensic-grade facial comparison?

Full forensic reports with detailed similarity scoring. Results in seconds.

Run My First Search