CaraComp
CaraComp
Forensic-Grade AI Face Recognition for:
Get Started7-day refund guarantee**
digital-forensicsBy Cara Candelario

Deepfake Audio: Inside the Voice Cloning Behind Arup's $25M Loss

The $25M Deepfake Used Three AI Layers at Once — How Each One Fooled a Human

Quick answer

How does deepfake CFO fraud work and why do employees fall for it?

Deepfake CFO fraud works by cloning an executive's face and voice from public footage, then staging a video call with several fake colleagues. Employees comply because a group of familiar senior figures agreeing creates pressure, even when something looks slightly off. Checking the request through a separate, previously known channel is the strongest defense.

The employee noticed something was off. The CFO looked slightly strange on the video call. The expressions felt slightly delayed, the movements slightly mechanical. He noticed, and then he transferred $25 million anyway.

That detail is the whole story. Not the technology. Not the AI. The fact that a trained professional saw the artifacts, registered the wrongness, and proceeded regardless, because five other senior executives on the same call all agreed the transfer should happen. All of them were fake. Every single participant on that call, other than the victim, was an AI-generated impostor.

TL;DR

The Arup $25M deepfake attack succeeded not because the technology was undetectable, but because three simultaneous AI systems, facial mapping, voice cloning, and behavior synthesis, created just enough convincing pressure to override a human's visual suspicion.

This is what investigators need to understand: modern deepfake fraud doesn't need to be perfect. It needs to be just convincing enough to tip the balance while social pressure does the rest. The technical pipeline behind that call is more specific, and more teachable, than most coverage suggests.


Facial Mapping in Arup Fraud: The 68-Point Skeleton

Before a single frame of deepfake video gets rendered, the attack begins with source material. In the Arup case, the World Economic Forum reported that attackers scraped LinkedIn profiles, press conference recordings, and YouTube appearances, all publicly available, no hacking required. The raw footage becomes training data.

What the algorithm actually does with that footage is more precise than "learns what the person looks like." It builds a geometric skeleton. Facial landmark detection maps 68 anatomical anchor points across the target face: the outer corners of each eye, the peaks of the cupid's bow, the attachment points of the earlobes, the widest points of the nostrils, the edges of the jawline. These 68 coordinates become a mathematical description of that face's geometry, the proportional distances, the angles, the depths.

That skeleton then drives 3D face reconstruction. According to a comprehensive technical review of face deepfakes published on arXiv, the process involves detecting 2D landmarks in both the source and target faces, computing the 3D pose that accounts for viewpoint and expression, segmenting the face from its background using a pre-trained neural network, and then warping the source face onto the target using alignment calculated from those 3D poses. The result: the attacker's head movements drive the CFO's face.

Here's the forensic implication investigators rarely hear: extreme angles break the warping algorithm. When a face tilts beyond roughly 40 degrees, the geometric alignment between the real skull and the reconstructed surface starts to fail. Jawlines drift. Ear geometry becomes inconsistent. The hairline flickers. These are not subtle artifacts, they're structural failures that frame-by-frame analysis can surface. The victim's "something looks off" feeling was almost certainly responding to exactly this.


Layer Two: Three Seconds of Deepfake Audio

The voice track runs on a completely separate system, and the training data requirements are shockingly low. This article is part of a series, start with Deepfake Calls Surge As Governments Bet On Biometr.

3 sec
of audio required to produce an 85% voice match, according to McAfee research
Source: McAfee AI Research

McAfee's research found that just 3 seconds of reference audio produces an 85% voice match. High-fidelity cloning, according to ThreatLocker, requires roughly 30 seconds of clean audio. MIT researchers demonstrated high-quality speech generation from only 15 seconds of training material. The CFO of a major international engineering firm had given dozens of recorded presentations. The attackers had more than enough.

Voice cloning works by extracting a speaker's acoustic fingerprint, the specific harmonic patterns, resonance characteristics, and prosodic rhythms that make a voice recognizable, and encoding it into a neural model. New text is then synthesized in that voice in real time. The output isn't a recording of the real person; it's a mathematical prediction of what that person would sound like saying words they never said.

The result, layered over the facial deepfake, creates a perceptual double-bind. The viewer's brain is simultaneously processing visual and auditory signals that both seem to match a known person. Even when one channel feels slightly wrong, the other channel pushes back toward trust. That's not a technical coincidence, it's the attack strategy.


Layer Three: The Behavior Problem (and Why Pre-Rendering Matters)

Here's a detail that doesn't get nearly enough attention: the deepfake in the Arup case almost certainly wasn't being generated dynamically in real time. SoftwareSeni's technical breakdown of the deepfake pipeline notes that typical tools require approximately 30 minutes of processing to generate a few sentences of convincing video. Real-time synthesis at that quality level, in 2024, was not commercially available at consumer price points.

So what the attackers actually did was closer to staging a play than running an AI model. They wrote a script. They pre-rendered a library of video clips, the CFO explaining the transaction, responding to expected questions, showing agreement, and then played those clips back in sequence during the call while a human operator controlled the timing. The "live" video call was a controlled playback, not a live generation. Previously in this series: 64 Deepfake Laws Passed And Investigators Still Ca.

Think of it like an elaborate puppetry show where the puppeteer has studied their target so thoroughly, voice recordings, video appearances, facial mannerisms, that they can perform a convincing one-person play. The audience isn't looking for puppet strings. They're listening for tone, watching for familiar expressions, and trusting the context. A good puppeteer with rehearsed lines and pre-recorded audio only needs to fool the audience for ten to fifteen minutes. That's exactly the window deepfake attackers require.

"The realistic visuals and audio, combined with the presence of multiple seemingly familiar senior figures discussing the transaction, ultimately convinced the employee of the request's legitimacy." Security Boulevard, analysis of the Arup deepfake attack

Behavioral artifacts are where trained examiners still have purchase. Eyeblink research has shown that real video contains periodic, biologically consistent blinking patterns, deepfakes frequently miss the timing, producing either unnaturally regular blinks or long stretches without any. Micro-expressions, gaze tracking, and the subtle asymmetry of genuine emotional responses all carry signals that pre-rendered deepfakes struggle to replicate consistently across an extended conversation. At CaraComp, frame-level facial comparison against baseline reference footage, not a confidence score, but a methodical landmark-by-landmark forensic comparison, remains one of the few approaches that surfaces these inconsistencies reliably.


Trusted by Investigators Worldwide
Run Forensic-Grade Comparisons in Seconds
Detailed facial comparison reports. Results in seconds.
Get Started
7-day refund guarantee**

Deepfake CFO Fraud: The Misconception That Spreads

Most people who hear about the Arup case walk away with a specific wrong conclusion: that the deepfake was so perfect, it was indistinguishable. This belief is both understandable and genuinely dangerous.

It's understandable because every news headline emphasizes how realistic modern deepfakes are becoming. "Indistinguishable." "Undetectable." If the story is that a trained professional at a major firm got fooled, surely the technology must be extraordinary.

But the victim explicitly said the CFO looked "a little off." He saw the artifacts. The deepfake was detectable, and it won anyway. That's a completely different threat model, and it demands a completely different response.

The attack succeeded through social engineering, not technical perfection. Five "executives" on the same call all confirming the same high-pressure transaction is not normal behavior, it's a manufactured consensus designed to override individual doubt. The deepfake only needed to be convincing enough to prevent the victim from stopping the call and making a phone call to a known number. That's a much lower bar than "undetectable."

What You Just Learned

  • 🧠 The three-layer architectureFacial mapping (68 landmarks), voice cloning (as little as 3 seconds of audio), and behavioral pre-rendering work simultaneously, not sequentially
  • 🔬 Extreme angles break the geometryThe facial warping algorithm fails at sharp head angles, producing jaw, ear, and hairline artifacts that frame-level analysis can detect
  • 🎭 It was staged, not livePre-rendered clip libraries played back in sequence, not real-time AI generation, which means the "conversation" was scripted and the response range was limited Up next: The 25M Deepfake Used Three Ai Layers At Once How .
  • 💡 The deepfake didn't need to be perfectSocial pressure from multiple fake "executives" overrode the victim's visual suspicion; the technology only needed to delay skepticism for 10 minutes

The One Verification Step That Changes Everything

The economics here are worth sitting with for a moment. Deepak Gupta's detailed analysis of the Arup case estimates the attack cost less than $10,000 to execute against a $25 million target, a roughly 2,500-to-1 return ratio. Even if only one in a hundred attempts succeeds, the math works overwhelmingly in the attacker's favor. This isn't a complex nation-state capability anymore. It's commercially available technology with extraordinary ROI.

Which means the response can't be "train people to spot deepfakes." That's a losing arms race against a system specifically designed to defeat human visual judgment. The response has to be procedural friction, and specifically, friction that exploits the one thing deepfake calls cannot easily defeat: multiple independent communication channels.

A video call request for an urgent financial transfer should trigger a callback to a number stored in your contact system before the call happened, not a number provided during the call. An email confirmation to an address in your existing directory. A second approver reached through a separate channel. These steps feel inefficient. That's the point. The Arup attack worked because efficiency was prioritized over verification. Sometimes friction is security.

Key Takeaway

Modern deepfakes don't need to fool expert forensic analysis, they only need to hold up for ten minutes under social pressure. The defense isn't better visual detection. It's verifying high-stakes requests through a second channel that was established before the suspicious call ever started.

If a key witness or claimant only ever appears to you on video calls, there's one question worth asking yourself right now: do you have a pre-established, out-of-band way to confirm their identity that doesn't run through the same session you're already in? If the answer is no, and for most investigators, it currently is, you're relying entirely on a signal that a $10,000 AI system was specifically built to spoof.

The victim in Hong Kong noticed something was wrong. He just had no protocol for what to do when the CFO looked slightly off but five other executives all said everything was fine. That's the gap. Not the technology. The gap is the absence of a procedure that treats video-only verification as insufficient for high-stakes decisions, because, frame by frame and landmark by landmark, it increasingly is.

Why Audio Deepfakes Are the Harder Problem to Detect

Of the three layers used in the Arup attack, deepfake audio is arguably the one investigators are least equipped to examine. Facial mapping leaves geometric traces. Behavioral pre-rendering leaves timing gaps. But audio deepfakes can be judged almost entirely by ear, and the human ear is not a reliable forensic tool. A convincing audio deepfake only needs to pass the listener's gut check for a few minutes, not survive a spectrogram review.

That asymmetry matters for how investigators should prioritize their attention. When a case involves deepfake audio, the voice sample itself should be treated as evidence and preserved in its original file format before any transcription or summary replaces it. Compression, re-recording, or playback through a phone speaker destroys the acoustic detail that a specialist would need to compare against genuine reference recordings. Investigators who only keep notes about what a caller said, rather than the audio itself, have thrown away the one thing that could later confirm or rule out a synthetic voice.

How Voice Cloning Turns Recordings Into a Weapon

Voice cloning does not require a hacker to break into anything. It requires access to audio the target has already made public, and most executives, public officials, and expert witnesses have far more of that than they realize. Every recorded webinar, deposition, podcast interview, or conference panel becomes a potential training set for someone building deepfake audio of that person's voice.

This is why the volume of clean audio matters more than its source. A single podcast appearance can supply more usable seconds of a person's speech than an entire year of casual phone calls, because studio or webinar audio is typically clearer and freer of background noise. Investigators assessing whether someone is a plausible deepfake audio target should ask a simple question: how much clean, publicly available audio of this person's voice exists online right now? If the answer is more than a minute or two, the technical barrier to cloning that voice is already gone.

Detecting Audio Deepfakes: What Actually Works Right Now

Detection of audio deepfakes is improving, but it still depends heavily on access to the original audio file rather than a description of a call. Specialists look for unnatural pauses between words, a flatness in emotional inflection during moments that should carry stress or urgency, and breathing patterns that don't match natural human speech rhythms. None of these checks can happen after the fact if the audio itself was never saved.

Detecting deepfakes in audio also benefits from comparing a suspect recording against multiple genuine samples of the same person speaking in different emotional states, calm, frustrated, urgent, because a cloned voice model trained mostly on calm presentation audio often fails to reproduce authentic stress patterns convincingly. A synthetic speech sample that sounds perfectly composed during a supposedly urgent, high-pressure financial request is itself a red flag worth documenting. This single mismatch has been enough, in other reported cases, to make a callback the difference between a stopped fraud and a completed one.

Cloned audio and cloned voice recordings both degrade in specific, checkable ways when they are stretched beyond the length of their training data. A voice clone built from three seconds to thirty seconds of source material performs best on short, predictable phrases and starts to strain on longer, more spontaneous exchanges. If a caller's voice sounds fluent and natural for scripted-sounding statements but oddly generic or hesitant when asked an unexpected follow-up question, that shift is worth noting. It suggests the speech was generated from a fixed script rather than produced by a person improvising in real time.

Fake audio and ai-generated audio built for fraud typically optimize for a narrow purpose: sounding convincing for a short, high-stakes exchange rather than holding up under sustained, unpredictable conversation. This is a practical weakness investigators can exploit without needing a lab. Asking an unscripted question, requesting the speaker repeat an unusual phrase, or introducing an unexpected topic change can expose the limits of even a well-made deepfake audio model, because the system behind it was never built to improvise beyond its rehearsed material.

An audio deepfake detector, where available, works by analyzing the same kinds of acoustic fingerprints that voice cloning tools extract in the first place, harmonic patterns, resonance, and prosody, and flagging recordings whose patterns look mathematically generated rather than naturally produced. These tools are not yet a substitute for procedural verification, but they are a useful second opinion when an original audio file has been preserved. Investigators who understand both how audio deepfakes are built and how audio deepfake detection tools examine them are far better positioned to know which cases deserve that extra layer of technical review.

What a Voice Changer Cannot Fake Under Pressure

A voice changer and a true deepfake voice model solve different problems, and investigators should not treat them as the same threat. A simple voice changer shifts pitch or tone in real time but does not model a specific person's speech patterns, which means it cannot reproduce someone's actual cadence, word choice, or breathing habits under stress. A deepfake voice built from cloned audio is a much harder problem precisely because it borrows real vocal characteristics rather than distorting a generic input voice.

Where a voice changer tends to fail fastest is spontaneous, emotionally charged speech. Real fear, frustration, or urgency changes breathing, pacing, and word emphasis in ways that a pitch-shifted voice changer simply cannot replicate, because it was never trained on that person's speech at all. This gives investigators a fast, low-cost screening question: does the caller's emotional register under pressure match what is known about how this person actually talks when stressed? A mismatch here is worth flagging even before any technical audio deepfake detection tool gets involved.

Why Deepfake Scams Depend on Speed, Not Just Realism

Deepfake scams involving audio cues rarely give a target time to think, and that time pressure is doing as much work as the fake voice itself. A deepfake voice or fake voice used in a wire-transfer scam is typically deployed inside a narrow window, a call demanding action in minutes, not hours, because the longer a target has to verify, the more likely the synthetic speech is to be caught. Investigators reviewing a suspected deepfake scam should note how much time pressure the caller applied, since urgency is itself a documented pattern across these cases. Resemble AI and Pindrop are two names that come up often in coverage of both sides of this problem: one focused on voice generation technology, the other on voice authentication and fraud detection. Understanding that these are different categories of technology, built by different companies for different purposes, helps investigators ask sharper questions when a vendor pitches "AI detection" software or "deepfake models" without specifying which problem it actually solves.

Frequently asked questions

What is deepfake audio and how does it work?

Deepfake audio is a synthetic voice created by extracting a speaker's acoustic fingerprint, meaning the harmonic patterns, resonance characteristics, and prosodic rhythms that make a voice recognizable, and encoding that into a neural model. New text is then synthesized in that voice. The output isn't a recording of the real person, it's a mathematical prediction of what they would sound like saying words they never said.

How much audio is needed to clone a voice with deepfake audio tools?

McAfee research found that just 3 seconds of reference audio produces an 85% voice match. ThreatLocker reported that high-fidelity cloning requires roughly 30 seconds of clean audio, while MIT researchers demonstrated high-quality speech generation from only 15 seconds of training material. In the Arup case, the CFO had given dozens of recorded presentations, giving attackers more than enough material.

How did deepfake audio contribute to the Arup $25 million fraud?

Deepfake audio was layered over a facial deepfake so that the viewer's brain processed matching visual and auditory signals at once, pushing toward trust even when one channel felt slightly wrong. Combined with pre-rendered video clips of the CFO and multiple fake senior figures on the call, the cloned voice helped convince the employee the transaction request was legitimate.

Ready for forensic-grade facial comparison?

Full forensic reports with detailed similarity scoring. Results in seconds.

Run My First Search