CaraComp
CaraComp
Forensic-Grade AI Face Recognition for:
Get Started7-day refund guarantee**
digital-forensicsBy Cara Candelario

AI Deepfake Fraud Detection: Video, Detect Deepfakes & Risk Signals

A 0.78 Match Score on a Fake Face: How Facial Geometry Stops Deepfake Wire Scams
A video call face is analyzed by facial geometry software, illustrating how ai deepfake fraud detection flags synthetic identities in real time.

A finance employee at a multinational firm sat on a video call with what appeared to be the company's CFO, along with several colleagues. Everyone looked right. Everyone sounded right. The call ended, and $25.6 million moved out of the company's accounts. Every face on that call had been deepfaked. Not pre-recorded, generated live, in real time, mapped over real humans sitting in a scam center somewhere, reading from a script.

The employee didn't make a mistake, exactly. They did what humans have been wired to do for 200,000 years: they looked at a face, heard a voice, watched the lips move, and trusted it. The problem is that the technology for faking all three of those things now fits inside a laptop.

TL;DR

Deepfake video call scams fool victims by overwhelming human sensory trust, but facial comparison technology bypasses psychology entirely, measuring the geometry of a face against a reference photo and returning a single mathematical score that no emotional narrative can override.

Why Your Eyes Fail at Deepfake Wire Fraud Detection

Here's the uncomfortable truth about human perception: we don't verify faces. We recognize them. Those are completely different cognitive processes. Recognition is pattern-matching against memory, fast, automatic, and almost impossible to consciously override. Verification is a deliberate, systematic check against an external reference. Humans are exceptional at the first. We are genuinely terrible at the second.

Scammers know this. The current operational standard for live deepfake fraud isn't a fully synthetic face floating in a void, it's a real human being hired as an "AI model," sitting on camera, while face-swap software maps a different identity over their face in real time. Malwarebytes has documented scam operations running dozens of simultaneous deepfake video calls per day, with hired performers whose faces are swapped to match whatever fictional (or stolen) identity the victim is expecting to see.

Voice cloning runs on the same call. And here's the part that should make investigators pause: according to The Register, citing Interpol research, a convincing voice clone now requires as little as ten seconds of reference audio, the kind you'd find on any public LinkedIn video, earnings call recording, or Instagram story. Ten seconds. That's all.

So when a victim sees a face they recognize, hears a voice they recognize, and watches lips that roughly sync with words, their brain says "this is real" before their skepticism even wakes up. That's not gullibility. That's neuroscience. And it's exactly the gap that facial comparison technology exists to close.

1 in 4 This article is part of a series, start with Age Assurance Becomes The New Kyc And Your Next Ca.
American consumers received a deepfake voice call in the past 12 months
Source: Hiya consumer survey across 6 countries, ~12,000 respondents

What Facial Comparison Actually Measures (It's Not What You Think)

Most people imagine facial recognition as something like a very fast visual comparison, the algorithm "looks" at two photos the way you would, just faster. That's wrong in the most interesting possible way.

What a facial comparison system actually does is convert each face into a vector: a list of numbers, typically 128 of them, representing the geometric relationships between specific landmarks on the face. The distance between the inner corners of the eyes. The ratio of nose bridge length to jaw width. The angle from the outer eye corner to the mouth commissure. These aren't chosen arbitrarily, they're the measurements that remain most stable across lighting changes, aging, and expression variation.

Once you have two 128-number vectors, one from your reference photo of the real person, one from the video call frame, you calculate the Euclidean distance between them. Think of it like measuring the straight-line distance between two points in 128-dimensional space. Academic research on deep learning face recognition models, including the widely studied FaceNet architecture, establishes a clear principle: faces of the same person produce consistently small Euclidean distances; faces of different people, or synthetic variants, produce larger ones.

A practical threshold looks something like this: a distance below 0.6 suggests the same identity; above 0.6 starts raising flags. A deepfake face from a video call, even a convincing one, often lands at 0.75, 0.80, sometimes higher, because the synthesis process introduces subtle landmark inconsistencies that the human eye never catches but the math immediately surfaces. That 0.78 in our headline isn't hypothetical scenery. It's the kind of number that stops a wire transfer cold.

"Deepfake fraud is no longer a future threat, it's a present operational reality. AI is now being weaponized at scale to impersonate individuals in real time, combining synthetic voice, video, and social engineering into attacks that defeat traditional verification entirely." Cyber Magazine, reporting on AI anti-deepfake platforms
Trusted by Investigators Worldwide
Run Forensic-Grade Comparisons in Seconds
Detailed facial comparison reports. Results in seconds.
Get Started
7-day refund guarantee**

The Forensic Analogy Behind Deepfake Fraud Detection

Ai Deepfake Fraud Detection And The Identity Fraud Problem

Identity fraud used to mean a stolen document or a forged signature. Ai deepfake fraud detection exists because identity fraud now happens on a live video call, in real time, with a face and voice that pass every casual human check. The detection software built for this problem doesn't try to spot a "bad actor" by instinct, it measures geometry and returns a number, which is the only thing that scales across thousands of calls a day.

Detection Software Versus Human Judgment

Detection software wins where humans lose: consistency under pressure. A tired investigator on their fifteenth case of the day applies less scrutiny than one on their first. Detection doesn't get tired, doesn't get rushed, and doesn't get talked into skipping the check because the caller sounded convincing. That consistency is the entire value proposition of automating deepfake detection inside a fraud workflow.

Think of facial comparison like a fingerprint match conducted at video-call speed. A fingerprint examiner compares ridge patterns point by point against a reference card, painstaking, expert-dependent, and measured in hours. Facial comparison does the geometric equivalent in about 30 seconds. Both methods test one specific forensic question: does this pattern belong to this person?

The speed is what turns the tool from interesting into something you can actually rely on in a live case. A victim on a live call is operating under time pressure, urgency is baked into the scam design. "Wire the funds before the window closes." "Don't tell anyone yet, this is sensitive." Scammers construct an emotional environment that makes deliberate verification feel like obstruction. A 30-second facial comparison check injects a hard fact into that emotional moment, a number, a threshold, a match or no-match, and that number cannot be talked out of existence by a convincing performance. Previously in this series: Deepfakes Force New Identity Rules And Investigato.

This is where the tool's real value lives. Not in the lab, but in the gap between "the call just ended" and "the transfer is authorized." At CaraComp, the principle behind this kind of check is exactly what drives how investigators should think about facial verification: isolate the visual claim, test it independently of audio, and let the geometry answer the question that psychology can't.

Deepfake Detection Tools And The Reality Defender Api Approach

Some vendors package deepfake detection as an API call rather than a manual review step. A platform built on something like a reality defender api model ingests a video frame or audio clip and returns a probability score for synthetic content, similar in spirit to the Euclidean distance score described above but trained specifically to catch generative artifacts rather than identity mismatches. Teams that combine both, identity distance scoring and a dedicated deepfake detection api, get two independent signals instead of one, which is harder for a scammer to defeat with a single trick.


The Misconception About Deepfake Fraud That Costs $40K

The belief that kills investigations before they start: "If it looked and sounded real on a video call, it must have been a real person."

It's worth being generous about why people get this wrong. For most of human history, a synchronized face-and-voice on a live call was causally impossible to fake at scale. The brain's shortcut, "live video plus matching voice equals real person", was a perfectly reasonable heuristic for the world as it existed in 2015. The problem is that heuristics don't update automatically when the world changes. The technology moved; the mental model didn't.

Here's what people miss about how deepfake generation actually works: voice synthesis and facial synthesis are trained separately, on different datasets, using different models. They're stitched together in post-processing (or real-time processing), but they're not one unified system. Which means each component has its own failure modes, its own artifacts, its own mathematical fingerprint. The audio can be perfect while the facial geometry fails. That's the seam investigators can find.

Facial comparison deliberately ignores the audio. It ignores the performance quality. It ignores whether the lip sync looked good or whether the person knew details only the real executive would know. It asks exactly one question, does the geometric structure of this face match the reference?and answers it with a number. That single-variable isolation is the methodological move that makes the difference.

What You Just Learned

  • 🧠 Facial recognition converts faces to 128 numbersnot a visual comparison, but a geometric embedding that enables mathematical distance scoring between two identities
  • 🔬 Euclidean distance is the key metrica score below ~0.6 suggests the same person; a deepfake face often scores 0.75 or higher, exposing the synthetic inconsistency
  • 🎭 Live deepfake scams use hybrid operationsreal humans on camera with AI face-swap running in real time, not pre-recorded clips, which is why old "ask them to turn sideways" tests are failing Up next: A 0 78 Match Score On A Fake Face How Facial Geome.
  • 💡 Voice and face are synthesized separatelyeven when audio is perfect, facial geometry can fail a comparison check, giving investigators an independent verification channel
Key Takeaway

A deepfake face that convinces a human brain still has to pass a geometric test it wasn't designed to pass. Facial comparison doesn't trust the performance, it measures the math. A 30-second check between a video call frame and a reference photo produces a number. That number doesn't care how urgent the caller sounded.

Global deepfake fraud losses hit $350 million in a single quarter of 2025, according to reporting by Resemble.ai cited across multiple security publications. That figure exists almost entirely because the verification gap, the space between "this looks real" and "this has been confirmed real", is still being filled by human judgment instead of geometric measurement.

The $25.6 million Arup case. The $40,000 wire scenarios playing out in smaller firms every week. Every one of them shares the same architecture: a convincing performance, a time-pressured decision, and no independent check on the face making the request.

The aha moment isn't that the technology exists, it's that the comparison is faster than the wire transfer approval process. An investigator, a compliance officer, or a fraud analyst who can extract a frame from a suspicious call and run a facial comparison against a LinkedIn photo or HR file has, in 30 seconds, answered the question that a $40,000 loss proves humans cannot answer reliably on their own. The victim's fear can override their eyes. It cannot override a Euclidean distance of 0.78.

So here's the question worth sitting with: If a client called you right now and said "I just got off a video call with my CEO asking me to approve an urgent transfer", what would be your first concrete move to test whether that face was genuine?

Ai deepfake fraud is now a line item in fraud budgets, not a hypothetical risk on a slide deck. Every finance team that approves wire transfers over video call is running deepfake fraud exposure whether they've named it that way or not. The organizations getting ahead of this aren't relying on training staff to "spot the fake", they're building a checkpoint where a face gets measured against a reference before money moves.

Deepfake detection at the identity layer works differently than deepfake detection built to catch generative artifacts in the video stream itself. Identity-layer detection asks "is this the specific person it claims to be," using a reference photo and a distance score. Artifact-layer detection asks "does this content show signs of synthetic generation," using patterns learned from training on known fakes. A mature ai deepfake fraud program runs both, because a scammer who defeats one system doesn't automatically defeat the other.

Content moving through a live video call, video, voice, and the metadata around the call itself, carries more forensic signal than most teams currently use. Voice can be checked against a known voiceprint. Video can be checked frame by frame against a reference photo. The call's technical metadata can reveal virtual camera software or routing patterns inconsistent with the claimed caller's location. None of these checks require the investigator to trust their gut; each produces a discrete, defensible data point.

Risk teams that build deepfake detection into the wire approval workflow, rather than treating it as a forensic afterthought, catch fraud before the money leaves, not after. The math doesn't change once you've already sent the wire. It only helps if it runs before authorization, which means the detection step has to be as fast as the fraud itself.

Digital identity verification has always rested on the assumption that a face and a voice are hard to fake convincingly at scale. Ai deepfake fraud detection is the response to that assumption breaking down. The tools don't need to be perfect; they need to be faster and more consistent than a rushed human glancing at a screen under pressure, and on that bar, they already clear it.

Video call fraud detection tools increasingly borrow techniques from both camps: use state-of-the-art detection tools trained on generative adversarial networks (GANs) output to flag synthesis artifacts, while also running the identity distance check described earlier in this piece. Removes deepfake fraud risk from any single point of failure, because a scam has to beat several independent measurements instead of one skeptical employee.

Detection deepfake workflows are also getting cheaper to run. What once required a specialized lab now runs as an API call that returns a score in seconds, which means even a small firm processing a handful of wire requests a week can afford to use multiple types of checks, identity distance, artifact detection, and voiceprint comparison, on every high-value call before approving a transfer.

Fraud teams evaluating a new detection stack should ask vendors to show their fraud detection rate against known deepfake video samples, not just marketing claims about accuracy. A vendor that can walk through how their model handles a low-light video call frame, a compressed video feed, or a video call recorded through a virtual camera driver is showing you real engineering, not a demo reel. Ask for the false-positive rate too, because a detection system that flags too many real employees as fraud risks gets disabled by frustrated staff within a month.

Doppel monitors brand impersonation and synthetic media across social platforms, which matters because deepfake fraud rarely starts and ends with a single video call. A scam operation that builds a convincing deepfake video often also runs fake executive profiles, cloned company domains, and synthetic social posts to make the story checks together, the video call is the final step in a longer con, not the whole thing. Teams that only watch for deepfake video and ignore the supporting synthetic media miss the setup phase where the fraud could have been caught earlier and cheaper.

Vendors building detect deepfakes capability into identity platforms are also shipping sdks so engineering teams can embed a video comparison check directly inside an existing onboarding flow or wire approval tool, instead of asking staff to open a separate application. An sdk approach means the fraud detection step happens inside the same screen where the transfer gets approved, which removes the extra click that busy staff are most likely to skip under deadline pressure.

Deepfake technology keeps improving on both sides of this fight, the generation side and the detection side, which is why a program built around one static checklist goes stale fast. A detection model trained on last year's deepfake video samples can miss artifacts introduced by a newer generation pipeline, so risk teams need a vendor that retrains regularly against fresh synthetic media rather than shipping a model once and leaving it alone.

Authentication has traditionally meant a password, a one-time code, or a security question, none of which help when the fraud is a live video call impersonating someone the victim already trusts. Layering facial comparison and voiceprint authentication on top of traditional authentication gives a finance team a second, independent gate that doesn't depend on the caller knowing a shared secret, because a well-resourced scam operation can often obtain those secrets too.

Social engineering is the part of a deepfake scam that technology alone can't fully patch, because the scammer is manipulating urgency and trust, not just spoofing a face. Combining a facial or voice detection check with basic social engineering awareness, training staff to expect a verification step on every high-value request, no matter how convincing the caller sounds, closes the gap that pure technology leaves open.

Media coverage of deepfake fraud cases keeps citing bigger dollar figures because attackers are getting better at picking high-value targets, not because the underlying detection math has stopped working. A finance team that reads about the next multimillion-dollar case should treat it as a reminder to check whether their own wire approval process includes an independent video or voice check, not as evidence that detection is a lost cause.

Building this into daily practice doesn't require a large security budget. A firm can start with a documented policy: any wire request that arrives through a video call or voice call above a set dollar threshold gets a facial comparison or voiceprint check before approval, full stop, no exceptions for urgency. That single rule, enforced consistently, closes the exact gap that the $25.6 million case and the $40,000 cases both walked through.

Frequently asked questions

How does ai deepfake fraud detection actually work?

AI deepfake fraud detection converts a face into a vector of roughly 128 numbers representing geometric relationships between facial landmarks, like eye spacing or jaw-to-nose ratios. It then calculates the Euclidean distance between that vector and a reference photo's vector. A distance below 0.6 suggests the same identity, while distances of 0.75 or higher, common in deepfakes, raise flags.

Why can't people just spot deepfakes with their own eyes?

Human perception recognizes faces through fast, automatic pattern-matching against memory rather than deliberate verification against a reference. Scammers exploit this by hiring real people on camera and using face-swap software plus voice cloning, which can be built from just ten seconds of audio, so victims trust what looks and sounds familiar before skepticism kicks in.

How much money can be lost to a deepfake video call scam?

In one documented case, a finance employee at a multinational firm joined a video call where every participant, including someone appearing to be the CFO, was a live deepfake generated in real time over real humans in a scam center. The call ended with $25.6 million moved out of the company's accounts.

Ready for forensic-grade facial comparison?

Full forensic reports with detailed similarity scoring. Results in seconds.

Run My First Search