CaraComp
CaraComp
Forensic-Grade AI Face Recognition for:
Get Started7-day refund guarantee**
digital-forensics

How to Detect Deepfakes: Audio, Data & Digital Clues

Stop Watching the Face: 3 Places Deepfakes Quietly Fall Apart
A close-up of mismatched lip movement illustrates how to detect deepfakes by watching mouth timing, shadows, and edges.

Here's something that should change how you watch video forever: the most convincing part of a deepfake is the face. Which means if you're staring at someone's face trying to decide if a video is real, you're looking in exactly the wrong place.

TL;DR

A fake video can pass the "face looks real" test and still betray itself in three quieter places: the mouth's timing, the shadows, and the blurry edges around hair or glasses — and knowing this one shift in where you look is your best defense.

Deepfake technology (software that replaces or generates a person's face and voice in video) has gotten genuinely good at faces. That's where most of the computing power goes. But video isn't a single picture. It's thousands of pictures — frames — playing back-to-back thirty times a second. Keeping every tiny detail perfectly consistent across all those frames? That's where things quietly break down. And that's exactly what you can learn to catch.


Why the face is the easy part

Think about what a deepfake actually has to do. It doesn't just need to generate one good-looking face. It needs to generate a face that moves naturally, responds to light the way real skin does, lines up with audio that was recorded separately, and does all of that perfectly across every single frame of the video. Miss by a fraction of a second in one spot, and the whole thing starts to feel wrong — even if you can't immediately say why.

CaraComp DailyEP.123
3 stories · 2:45
Starts at 01:37 — this story
2:45

Watch this story, in under a minute

Plays right here · jumps to 01:37
In this episode

A new briefing every weekday — three stories, three minutes.

Subscribe on YouTube

The face itself? Relatively straightforward for modern AI. Skin texture, eye movement, general facial structure — the tools for faking those have been improving for years. But the relationships between things in the frame are much harder to fake. The relationship between mouth movement and sound. Between a light source and the shadow it should cast. Between the edge of someone's hair and the background behind them. Those are the seams.

"Deepfakes are no longer easy to dismiss as clumsy edits — they can imitate a person's face, voice, expressions and mannerisms well enough to lend authenticity to a false message." Bitdefender Hot for Security

Which is exactly why you need to know where to look. Because the cues are still there. They're just quieter than you'd expect.

Deepfake detection starts with knowing what to look for

Deepfake detection isn't a single test — it's a habit of checking several small things at once instead of trusting one overall impression. A deepfake can nail the face and still fail the mouth timing, the shadows, or the edges around hair and glasses. Once you know the short list of places deepfakes tend to slip, you can run through it in under a minute on any video that feels suspicious.

How to recognize deepfake video without special tools

You don't need software to recognize deepfake content most of the time. Slow the video down if your player allows it, and watch the mouth during closed-lip sounds, the shadow direction on the face versus the room, and the edges where hair meets background. Any one of these being slightly wrong is a clue; two or three wrong at once is a strong signal you're looking at a deepfake.

Identifying deepfakes by focusing on relationships, not faces

Identifying deepfakes gets much easier once you stop grading the face on its own and start grading how it relates to everything around it. Does the mouth match the audio? Does the shadow match the light in the room? Does the hairline stay sharp as the person turns their head? Identifying deepfakes this way works even as the underlying deepfake technology keeps improving, because it targets a structural weakness, not a current technical limit.

Deepfake audio has its own tells, separate from the video

Deepfake audio is worth checking on its own, apart from whether the mouth matches the sound. Listen for a flat, slightly metallic quality, unnatural pacing between words, or breathing sounds that don't quite land where a real speaker would pause. Deepfake audio often nails individual words while missing the small human noises — a swallow, a breath, a stumble — that real speech includes without effort.


Detecting deepfakes: the mouth's timing

Of the three places deepfakes slip up, the mouth is the most interesting — and the most specific. Here's why. This article is part of a series — start with That Too Perfect Video 4 Hidden Clues Its Fake.

When you speak, your lips don't just open and close randomly. Every sound you make has a precise shape that goes with it. Linguists call these shapes "visemes" (the visual version of a sound). The sounds M, B, and P? Those require your lips to press completely together. No exceptions. If someone on screen appears to say "maybe" but their lips never fully close during the M or B sounds, something is wrong. The audio and the mouth geometry don't match.

This is a real detection target that researchers use. According to research published on ResearchGate, a detection framework called LIPINC tracks exactly these mouth-region inconsistencies across adjacent video frames — and catches what the human eye misses in real time.

The timing threshold is startlingly small. Mismatches that last as little as 50 to 100 milliseconds — that's a tenth of a second or less — can signal manipulation, even when the lip-sync looks convincing at normal speed. You will not catch this by watching normally. But you might notice that something feels slightly "off" about how someone's mouth moves, even if you can't name it. That feeling is worth paying attention to.

95%+
detection accuracy when analyzing mouth timing and phoneme-to-viseme alignment across frames
Source: Audio-visual temporal inconsistency research, via DuckDuckGoose AI

Compare that to how well humans do at spotting deepfakes by visual inspection alone: roughly 60 to 70 percent accuracy. That gap — between what your eyes catch and what frame-by-frame analysis catches — is the whole story. The forgery doesn't live in the pixels. It lives in the motion between frames.

Where gaze patterns often break first

Gaze patterns are another quiet tell, and they sit right next to the mouth as a place worth checking. Real eyes drift, refocus, and blink in small irregular bursts; generated eyes in a deepfake sometimes hold too steady or shift in a way that looks slightly mechanical. Watching gaze patterns for a few seconds, alongside the mouth timing, gives you a second independent check instead of relying on just one signal. Gaze patterns often break first around blinking rate and the tiny involuntary eye movements called saccades, which are difficult for generated video to reproduce naturally over many frames.


Where deepfakes slip: shadows and edges

Light has rules. If a light source is to someone's left, the shadow falls to their right. The brightness on their nose, their cheekbone, their chin — all of it follows those rules consistently. Real video gets this automatically, because light in the real world just works that way.

Generated video has to calculate all of that. And when a face is swapped onto a body filmed under different lighting, or when AI generates the face from scratch, the lighting math sometimes doesn't line up. The face might look a little too bright compared to the neck. A shadow might not quite match the angle of the light source in the room behind the person. These aren't obvious — but they're there.

Then there are the edges. The boundary between a person's hair and the background. The frame of someone's glasses where they meet skin. These transition zones are genuinely hard for AI to render cleanly across thousands of frames. Look for blurring, flickering, or a slightly "cut-and-pasted" quality right at those boundaries. Real video doesn't have that. The hair moves. The glasses shift slightly. The edges stay sharp and consistent. Faked video often can't keep up. Previously in this series: That Tattoo In The Photo Its Now Searchable But It Cant Prov.

Here's a good analogy for how all of this works together. Imagine a performer lip-syncing to a pre-recorded song on stage, trying to convince the audience she's singing live. From your seat, her mouth movements look roughly right. But if you're watching the exact moment the chorus hits and timing it against when her lips actually form the right shapes — you'll catch the lag. The song is real. Her voice is real. But the synchronization between them is off by just enough. That's the deepfake problem in a single image: not the face, not the audio, but the relationship between moving parts that refuses to stay perfectly in sync across hundreds of frames.

Deepfake skin often appears too smooth

One more physical tell sits right at the surface: deepfake skin often appears too smooth in a way real skin rarely is. Pores, faint texture, and small blemishes on the forehead and cheeks are hard for generated video to reproduce consistently, so a face that looks slightly airbrushed across the forehead and cheeks, even in poor lighting, is worth a second look. Combine that with checking whether the skin holds a clear spatial relationship to the jawline and neck as the head turns. On the opposite end, skin that looks too wrinkly or textured in patches that don't match the person's apparent age is also worth noting, since generated skin sometimes overcorrects for smoothness by adding uneven texture instead.

Visual inconsistencies add up faster than any single clue

No single clue proves a video is fake, but visual inconsistencies rarely travel alone. A mismatched shadow, a slightly blurred hairline, and skin that looks too smooth or too wrinkly in the same clip are visual inconsistencies that reinforce each other. When two or three visual inconsistencies show up together, treat that as a much stronger signal than any one oddity on its own.


Trusted by Investigators Worldwide
Run Forensic-Grade Comparisons in Seconds
Detailed facial comparison reports. Results in seconds.
Get Started
7-day refund guarantee**

The misconception that makes deepfakes work

Here's why people get fooled, and it's not because they're naive. It's because of how human brains are built.

We evolved to recognize faces. It's one of the oldest, most automatic things our brains do — so automatic that there's an entire region of the brain (the fusiform face area) dedicated specifically to it. When we see a face that looks like someone we know, trust kicks in before skepticism gets a chance. The face is the first thing you look at. It's the thing that feels most "real" or "fake" to your gut.

Deepfake creators know this. The face is where effort goes, because the face is where your attention goes. What they're banking on is that you won't look at the mouth timing, the shadow consistency, or the hair edges — because nobody naturally does. You're watching the eyes. You're reading expression. You're doing the thing your brain was designed to do.

The correction isn't "stop trusting your instincts." It's "add one more question." Instead of asking does this face look real? — ask does the whole scene behave naturally? The face and the audio. The light source and the shadows it should cast. The edges and whether they hold steady as the person moves. That's a different kind of watching. And it's one you can actually learn.

What You Just Learned

  • 👄 The mouth is the most specific tell — sounds like M, B, and P require complete lip closure; if those closures don't happen, the audio and the face were generated separately
  • 🔦 Shadows and lighting have rules — a fake face placed on a real body often breaks those rules in subtle ways the eye can sense but not always name
  • ✂️ Edges give it away — hair boundaries, glasses frames, and skin transitions are genuinely hard to render consistently across thousands of frames
  • 🧠 The forgery lives between frames, not in them — any single frame of a deepfake can look perfect; the problem shows up in motion and timing across the sequence

At CaraComp, we spend a lot of time thinking about exactly this — how AI systems analyze faces across time, not just in snapshots. The same principles that make facial recognition reliable (tracking consistency frame-to-frame, checking whether lighting and geometry stay coherent) are the ones that expose deepfakes. Real faces are consistent in ways that are surprisingly hard to fake at scale. Up next: Deepfake Detection Trust Infrastructure Three Layers.


What to actually do before you share or react

You don't need to become a video forensics expert. You need one habit: slow down before you believe something urgent.

If a video shows someone you know — a family member, a public figure, a coworker — saying something surprising or asking for something, here's where to aim your attention. Watch the mouth, not the face. Look for moments where the lip movement feels slightly ahead or behind the words, especially on sounds that should close the lips completely. Check whether the lighting on the face matches the lighting in the room. Look at the hair edges and glasses frames as the person moves — do they stay sharp, or do they shimmer and blur?

And if the video is paired with urgency — "act now," "don't tell anyone," "send this today" — that pressure is itself a red flag. The US Federal Trade Commission recommends a specific response when you get an unexpected message from someone you know: hang up, and call them back on a number you already have saved. Don't use contact information from the suspicious message itself. That one step — independent verification — protects you regardless of how convincing the video looks.

Key Takeaway

A convincing face is not proof a video is real — it's actually a reason to look harder at everything else. Check the mouth timing on sounds that need lip closure, watch whether shadows match the light source, and look at the edges around hair and glasses. The deepfake is almost never hiding in the face. It's hiding in the gaps between things.

Here's the thing that should stick with you: the better deepfake technology gets at making faces look real, the more useful it becomes to look somewhere else entirely. The face is the performance. The mouth timing, the shadows, the edges — those are the stage machinery. And stage machinery is always harder to hide than the actor standing in front of it.

So the next time a surprising video lands in your feed or your inbox — before you feel the pull to react, forward, or believe — ask yourself one question: does the whole scene behave naturally? Not just the face. The whole scene. That single shift in how you watch might be the most useful thing you read this week.

Most people trying to catch a deepfake reach straight for detection tools, but understanding what those tools actually check is more useful than any single app. Automated detection tools generally look for the same things this article describes: mismatched mouth timing, inconsistent shadows, and rough edges around hair and glasses, just measured in numbers instead of by eye. Knowing what the software is checking makes you a better judge of whether its verdict makes sense for a given video.

Digital forensics is the broader field that deepfake detection sits inside, and it treats video the same way it treats any other piece of potential evidence. Digital forensics experts look at compression patterns, metadata, and frame-level inconsistencies that never show up to the naked eye. You don't need that level of rigor for everyday scrolling, but knowing digital forensics exists — and that it can be called on for serious cases — is useful context.

Media forensics narrows that same idea specifically to photos and video, which is exactly the territory deepfakes live in. A media forensics review typically checks lighting consistency, edge quality, and audio-visual sync — the same three areas this article walks through, just done with specialized software instead of a careful eye. If a video matters enough to verify formally, media forensics is the field that does it.

Deepfake detectors built into apps and browser extensions are becoming more common, and most work by scoring the same red flags a careful viewer would notice: mouth-audio mismatch, shadow inconsistency, and edge blur. No deepfake detector is perfect, and a low-confidence score doesn't guarantee a video is real. Treat deepfake detectors as one more data point, not a final verdict, especially when the stakes are high.

When you're scanning a video for red flags, it helps to have a short mental list rather than a vague sense of unease. The core red flags are: lips not fully closing on M, B, and P sounds, shadows that don't match the visible light source, and blurry or flickering edges around hair and glasses. Three or four red flags together are far more convincing than any single one on its own.

Pay attention to how the media in question reached you, too — not just what's inside the video itself. A video shared through an unfamiliar link, forwarded many times, or stripped of its original source is worth more scrutiny than one posted directly by a known account. The path media takes before it reaches you is itself a clue, separate from anything visible in the frame.

None of this requires special training. It requires slowing down, checking the mouth, the shadows, the edges, the skin, and the gaze, and treating any single red flag as a reason to look for a second one before you believe or share what you're watching.

Learning how to detect deepfakes is also a security skill, not just a media-literacy trick, because the same techniques that fool casual viewers are used in scams that target bank accounts and workplace trust. A basic security habit — checking mouth timing, shadows, and edges before acting on a video — closes off one of the easiest ways scammers currently get people to move money or share information. Treat deepfake awareness the way you'd treat any other security practice: something you apply by default, not just when a video already feels suspicious.

Security teams inside companies are increasingly training employees to recognize deepfakes for the same reason banks train tellers to spot forged checks. A short security briefing on mouth timing, lighting mismatches, and edge blur can stop a fraudulent video call before it results in a wire transfer or a leaked password. If your workplace hasn't covered this yet, the red flags in this article are enough to bring to a security discussion.

Personal security also benefits from knowing how to detect deepfakes, especially as voice and video scams targeting families become more common. A grandparent scam or a fake emergency call is harder to pull off if the target already knows to check for mismatched mouth timing or odd lighting before reacting. That small layer of security awareness turns a moment of panic into a moment of verification.

The content of a video matters less than whether that content holds together under a second look. A deepfake can carry any content — a request for money, a political claim, a piece of gossip — but the underlying tells stay the same regardless of what the content is about. Judging content by its emotional pull alone is exactly what deepfake creators are counting on.

Media literacy and security awareness overlap heavily here, because both come down to pausing before you trust what a screen shows you. The same media habits that help you catch a deepfake — checking mouth timing, shadows, and edges — also help with other kinds of manipulated media, like selectively edited clips or misleading captions. Treating every piece of surprising media with the same quick checklist keeps you consistent instead of only skeptical when something already feels off.

Deepfake video and deepfake audio usually travel together, since most convincing fakes pair a generated face with generated or cloned speech. Checking deepfake video clues like mouth timing and edges, then separately checking deepfake audio for flat tone or odd pacing, gives you two independent tests instead of one combined impression. A video that passes one check but fails the other is still worth treating as suspicious.

Data about where a video originated is often more revealing than the video itself. Basic data like upload date, original account, and whether the file has been re-encoded multiple times can be checked in seconds and often tells you more than a frame-by-frame review. When that data is missing or scrubbed, treat the missing data itself as a signal worth factoring in.

Voice is worth judging on its own terms, separate from how it lines up with the mouth on screen. A cloned voice can sound right on individual words while missing the natural rhythm, breathing, and small imperfections of a real voice recorded live. If the voice sounds slightly too smooth or evenly paced for the emotion being expressed, that mismatch between voice and context is worth noticing.

Audio deserves the same frame-by-frame patience that video gets, because audio manipulation often hides in short gaps a casual listen skips over. Background audio that cuts out unnaturally, or audio that doesn't shift at all as the camera angle or room changes, is a sign the audio track may not have been recorded with the video. Trust your ear the same way you'd trust your eye on shadows and edges.

Threat actors who use deepfakes for fraud are counting on speed, not sophistication, because most people don't slow down enough to apply any of these checks. Understanding the threat in practical terms — a scammer wants you to act before you verify — is often more protective than knowing every technical detail of how the fake was made. Treating every urgent, unusual video request as a potential threat by default costs you nothing and protects you a great deal.

Cybersecurity teams treat deepfakes as one more entry on a growing list of social-engineering threats, alongside phishing emails and spoofed phone numbers. The same cybersecurity instinct that tells you to verify a suspicious email — check the source, don't act on urgency alone — applies directly to suspicious video and audio. Basic cybersecurity habits at home protect you just as much as they protect a company network.

Frequently asked questions

How to detect deepfakes without staring at the face?

The trick is looking away from the face, since that's the part deepfake software renders most convincingly. Instead, watch the mouth's timing against the audio, check shadows for natural light response, and study blurry edges around hair or glasses. These quieter details are harder for the technology to keep consistent across every frame.

Why does mouth timing help detect deepfakes?

A deepfake has to sync a generated face with audio recorded separately, and matching that perfectly across thousands of frames playing thirty times a second is difficult. When the mouth's timing is off by even a fraction of a second, the video starts feeling wrong, even if a viewer can't immediately explain why.

What visual clues reveal a deepfake video?

Shadows and edges are where deepfakes slip. Because video is made of thousands of frames, keeping tiny details like how shadows fall or how edges around hair and glasses stay sharp is hard to maintain consistently. Checking these blurry or inconsistent spots is more useful than judging whether the face itself looks real.

Ready for forensic-grade facial comparison?

Full forensic reports with detailed similarity scoring. Results in seconds.

Run My First Search