CaraComp
CaraComp
Forensic-Grade AI Face Recognition for:
Get Started7-day refund guarantee**
digital-forensicsBy Cara Candelario

Face Match Percentage: Two Faces, One Score, Four Checks First

A Facial Recognition 'Match' Isn't Evidence Until It Survives These 4 Hidden Steps
An analyst reviews facial comparison results, illustrating how software for facial recognition requires human verification before use as evidence.

Quick answer

How accurate is facial recognition as evidence in an investigation?

Facial recognition output is a lead to weigh with other evidence, not a conclusion. Accuracy depends heavily on the input image. Low resolution, angled faces, poor exposure and blocked features can degrade results while the score looks normal. A trained reviewer should check quality, threshold and the images side by side.

Here's a number that should stop you cold: a 50-percentage-point drop in accuracy can happen before the algorithm even gets a fair shot, simply because the inter-eye pixel distance in the probe image falls below 24 pixels. That's a resolution level you'll encounter constantly in real operational footage. The algorithm doesn't warn you. The confidence score doesn't flag it. The system just quietly becomes half as reliable as the benchmark chart promised, and the result lands on your screen looking exactly the same as it would if the image quality were perfect.

TL;DR

A facial comparison result is not a conclusion, it's a signal that needs to survive four distinct human and technical filters before it belongs anywhere near a report, a client, or a courtroom.

This is the gap between what people think facial recognition does and what it actually does. Most people's mental model is: upload photo, algorithm compares, system says "match," case closed. That model is approximately as accurate as thinking your GPS knows where you are because it's connected to satellites, technically true, but missing about six layers of engineering that determine whether the thing is actually right.

In serious investigations, the algorithm's output is where the story starts, not where it ends. And understanding the four hidden steps between a raw similarity score and a result you can put your name on is, increasingly, the difference between rigorous professional practice and expensive mistakes.


Facial Recognition Evidence: Quality Assessment Fundamentals

Before any comparison algorithm runs, something more fundamental has to happen: the input image has to be assessed on its own terms. Not "is this face similar to that face?" but "is this image even workable?"

CaraComp DailyEP.2
3 stories · 4:26
Starts at 02:38 — this story
4:26

A new briefing every weekday — three stories, three minutes.

Subscribe on YouTube

Research published in the International Journal of Legal Medicine found a direct, measurable relationship between image quality scores and match outcomes, high quality scores correlated with correct matches, low quality scores correlated with incorrect matches, and critically, high exposure was linked to false negatives while low exposure was linked to false positives. These aren't edge cases. They're the physics of how light and resolution interact with the feature-extraction process. This article is part of a series, start with China Made Creating A Deepfake The Crime Not Sharing It U S .

What does "image quality" actually mean here? It's not aesthetics. It's a set of measurable parameters: resolution (measured in inter-eye pixel distance), pose angle (algorithms trained on frontal faces start losing accuracy meaningfully when the face rotates beyond 30 degrees), illumination uniformity, occlusion percentage, and compression artifacts from CCTV encoding. Pose angles beyond 30 degrees can reduce match confidence scores by 30-40% even on top-performing algorithms, yet the score appears on screen as if that degradation never happened.

This is why experienced investigators treat the quality check as a first-line gate, not an afterthought. If the probe image fails a quality assessment, the downstream comparison score is not just imprecise, it's potentially misleading in ways that aren't self-announcing.

50pts
accuracy drop observed when inter-eye pixel distance falls below 24 pixels, a resolution level common in operational surveillance footage
Source: CaraComp operational accuracy research

Facial Recognition Accuracy: What Confidence Scores Don't Reveal

Let's say the image passes quality checks. The algorithm runs. A score appears, say, 0.94. Most people read this as "94% certain this is a match." That reading is understandable. It's also wrong, and understanding exactly why it's wrong is the cognitive shift that separates careful practitioners from overconfident ones.

A recognition confidence score describes the similarity between two templates extracted from the probe and reference images. It does not describe the probability that the two images show the same person. Those are different questions. And here's the kicker: the score is also shaped by image quality itself. A lower score can indicate poor quality images rather than less similarity between the people pictured. The number you're looking at conflates two separate things, actual dissimilarity and extraction degradation, into a single value, without telling you which one is driving it.

Think of it like a thermometer that changes its measurement scale based on ambient conditions. If the temperature reads 98.6°F, you don't know whether that person is healthy or whether the thermometer is running cold because it's been sitting in a drafty room. You need to know the instrument's state to interpret the reading. Same principle here.

Why do people get this wrong? Because numbers that look like percentages trigger a deeply ingrained mental shortcut, we treat them as probability statements. An algorithm that returns "0.94" is doing something technically precise, but not epistemically complete. The precision of the output implies confidence that the process doesn't actually deliver on its own. Previously in this series: Facial Recognition Isnt On Trial Your Explanation Is.

"The conversation cannot stop at whether the system performs well technically; organisations also need to consider how the technology is governed, how data is stored, who has access to it and whether its use can be clearly justified." Startups Magazine

Trusted by Investigators Worldwide
Run Forensic-Grade Comparisons in Seconds
Detailed facial comparison reports. Results in seconds.
Get Started
7-day refund guarantee**

Threshold Risk: Where Technical Settings Meet Legal Decisions

Here's where it gets genuinely interesting, and where most people's understanding of facial recognition has a blind spot the size of a barn door.

Every facial comparison system operates against a threshold: a cutoff score below which the system doesn't return a match candidate. This threshold appears to be a technical parameter. It is not. It is a risk decision, made by human beings, about which kind of error matters more in your specific context.

Set the threshold too high, and you'll miss genuine matches, legitimate subjects who should appear in your candidate list won't surface because their image quality dragged the score below the cutoff. Set it too low, and you'll be wading through low-quality candidates that waste investigator time and create false leads. As the Bipartisan Policy Center has documented through NIST benchmark data, at least six of the most accurate identification algorithms had higher false-positive rates for one demographic group at one threshold but lower false-positive rates at a different threshold. The same algorithm. Different threshold. Different demographic error profile. The "neutral" setting doesn't exist.

This is precisely why the National Institute of Standards and Technology measures algorithm performance at specifically defined false match rates, 0.001% and 0.0001%, rather than at a single universal threshold. The point is to make the tradeoff explicit. Every threshold is a statement about what your organization is willing to risk. Treating it as a default technical setting is the same as leaving that risk decision to whoever installed the software.


Step Four: The Human Review That Cannot Be Shortcut

The algorithm produces a ranked candidate list. The threshold determines what appears on that list. Then comes the step that no amount of algorithmic sophistication replaces: a trained investigator looks at the images, side by side, and makes a judgment.

This isn't a formality. Research in forensic facial comparison, including work published in ScienceDirect, establishes that forensic comparison systems need calibrated confidence measures, not raw scores, with the preferred approach being a score-based likelihood ratio that places the algorithm's output within a statistical framework. What this means practically is that translating a similarity score into evidentiary weight requires human expertise in reading the quality factors that shaped that score, not just the score itself. Up next: Law Enforcement Biometrics Facial Comparison Compliance.

At CaraComp, this is the step where the actual investigative value gets realized, where an analyst's understanding of what drove a particular score (image quality degradation? pose angle? partial occlusion?) transforms a raw number into a defensible conclusion. The algorithm's job is to prioritize. The analyst's job is to evaluate. Those are not the same job, and the analyst cannot skip the algorithm any more than the algorithm can replace the analyst.

What You Just Learned

  • 🧠 Image Quality Checkresolution, pose angle, exposure, and occlusion assessment before any comparison runs
  • 🔬 Algorithm Score & Thresholdsimilarity score generation against a threshold that reflects an explicit risk tolerance decision
  • 👁️ Investigator Visual Reviewtrained side-by-side comparison that contextualizes the score against quality factors
  • 💡 Report & Risk Decisioncalibrated confidence statement, not a binary match/no-match, suitable for a case file or legal context
Key Takeaway

A confidence score tells you how well the algorithm extracted features from a specific image. It does not tell you how confident you should be in the match. Those are different questions, and conflating them is where investigations, and boardroom risk decisions, go wrong.

NIST's own face verification testing data shows that recognition accuracy has improved dramatically since 2013, miss rates averaging 0.1% on high-performing algorithms in controlled conditions, with software from 2018 performing at least 20 times better than 2014 equivalents. But those numbers describe algorithm performance at its ceiling, on structured, high-quality datasets. Every piece of operational footage, every angled surveillance image, every low-light capture sits somewhere below that ceiling. The question is never "is this algorithm good?" The question is always "what's the quality of this specific input, and how far below peak performance is this specific comparison running?"

That question cannot be answered by looking at the score alone. Which is why the next time you see a facial match result, from any system, for any purpose, the first manual check isn't "is the score high enough?" It's "what was the quality of the image that produced this score?" One of those questions has an answer baked into the output. The other one requires you to go looking. That's exactly where rigorous practice lives.

When you get a "match" result in your work, what's the first manual check you run before you're willing to put your name on it? We'd genuinely like to know, the answer varies more across disciplines than most people expect.

Recognition Systems: How Buyers Actually Compare Options

Anyone evaluating recognition systems for a real operational deployment quickly learns that the vendor's headline accuracy number is the least useful part of the decision. What matters more is how the recognition platform behaves on your own image sources, your camera resolution, your typical lighting, your typical pose angles. A recognition platform that performs beautifully on studio-quality photos can still struggle badly on grainy CCTV frames, and no spec sheet will tell you that in advance.

Recognition Algorithms and the Limits of a Single Score

Recognition algorithms differ in how they extract features, how they weight quality signals, and how they respond to partial occlusion. Two different recognition algorithms can look at the same probe image and produce meaningfully different confidence scores, not because one is "wrong," but because they were trained on different data and tuned to different tradeoffs. That variability is exactly why a single number from a single algorithm should never be treated as the final word.

Face Recognition and Identity Verification Are Not the Same Task

Face recognition, matching an unknown face against a database of candidates, and identity verification, confirming that a person is who they claim to be, sound similar but carry different risks. Identity verification usually compares one live face to one reference photo, a narrower and generally easier problem than searching a large gallery. Treating the two as interchangeable is a common mistake that leads teams to apply the wrong threshold, the wrong review process, or the wrong level of scrutiny to the wrong task.

What to Ask Before Buying Facial Recognition Software

Before adopting facial recognition software for casework, ask the vendor how the system reports quality warnings alongside its scores, not just after a failure is discovered. Good facial recognition software surfaces resolution, pose, and occlusion problems at the moment of comparison, rather than burying them in a technical log nobody reads. If a vendor cannot explain how their matching process behaves on degraded images, that is itself useful information about how much manual review their tool will require.

Matching in Practice: Reading the Whole Picture, Not Just the Score

Matching two images well requires looking past the single output number and into the conditions that produced it. An investigator doing matching work learns to ask what the lighting was, what the camera angle was, and whether part of the face was blocked, before deciding how much weight the score deserves. That habit, treating matching as an investigation rather than a lookup, is the single biggest predictor of whether a team avoids costly misidentifications.

Face recognition software has become common enough that many organizations assume it works the same way in every context, but the underlying face recognition software still inherits every limitation described above: quality dependence, threshold tradeoffs, and the need for trained human judgment. Teams that treat face recognition software as a plug-and-play answer skip the review layer that actually protects them from bad outcomes. The software itself is only as reliable as the process wrapped around it.

Search remains one of the most misunderstood parts of the process, because a facial recognition search returns candidates ranked by similarity, not a verified identity. A search against a large gallery will always produce some result, even when the true match isn't present in the database at all, which is exactly why a returned candidate needs the same quality-and-threshold scrutiny as any other output. Treating a search result as an answer, rather than a lead, is where many of the costliest mistakes begin.

Privacy considerations sit alongside accuracy concerns, and they don't cancel each other out. An organization can run technically accurate facial recognition and still expose itself to serious privacy risk if it can't explain why a search was run, how long the images are retained, or who can access the results. Building privacy safeguards into the security process, access controls, retention limits, audit logs, is not separate from getting the technology right; it's part of the same responsibility.

Security teams adopting this technology should also budget time for the human review step described earlier in this article, because skipping it doesn't just create legal risk, it undermines the security value the technology was supposed to deliver in the first place. A fast, unreviewed match that turns out to be wrong is worse than no match at all, since it can send investigators down the wrong path while the real lead goes unexamined. The discipline of pairing technology with review is what makes security programs defensible rather than merely fast.

None of this means the technology itself is untrustworthy. It means the technology is a tool with known operating conditions, and like any tool, it performs well when used inside those conditions and poorly outside them. Software for facial recognition keeps improving on the metrics vendors publish, but published metrics describe best-case performance, not your case's performance. The gap between the two is exactly where the four-step process in this article does its work.

When someone asks why is my facial recognition not working inside an operational or investigative deployment, the honest answer is almost never "the algorithm is broken." A face match that returns nothing, or returns the wrong candidate, is usually the visible symptom of one of the four upstream steps failing quietly: a probe image that never should have cleared quality assessment, a threshold tuned for a different risk tolerance than the one your case actually needs, or a review step that got skipped under time pressure. Working backward through those four checks, rather than assuming the software itself is defective, is almost always the faster path to a real answer.

Face is the input every one of these systems depends on, and it is also the most variable input any comparison pipeline handles. The same face, photographed twice, can produce two different quality scores depending on lighting, angle, and camera resolution, which is exactly why a single face comparison should never be read as a stable, repeatable measurement. A face that looks clear to the human eye can still be a poor input for an algorithm if the pose angle or exposure falls outside the range the system was tuned on. Treating "the face was visible" and "the face was usable" as the same claim is one of the more common errors non-specialists make when reading a match result.

Not working is the description most people reach for when a facial comparison tool returns nothing or returns an obviously wrong candidate, but "not working" is rarely a binary state for these systems. More often, the system is working exactly as designed, extracting features and scoring similarity, while the input quality or the threshold setting quietly pulls the result away from what a human observer would expect. Diagnosing a facial comparison tool that appears to not be working starts with the same quality assessment described earlier: resolution, pose angle, exposure, and occlusion, checked in that order, before assuming the algorithm itself is at fault.

Camera conditions are frequently the real explanation behind a disappointing result. A camera positioned at a steep angle, or operating in low light, produces probe images that sit well outside the frontal, well-lit conditions most benchmark testing uses, and no amount of algorithmic sophistication fully compensates for that gap. Anyone troubleshooting a facial comparison result should look first at what the camera actually captured, resolution, angle, lighting, before assuming the comparison logic itself needs adjustment.

Reset is sometimes offered as a generic fix for systems that seem to be underperforming, but resetting a threshold or a configuration without first understanding why the current setting produced a given result just trades one blind spot for another. A reset threshold still needs to reflect an explicit decision about which kind of error your organization can tolerate; resetting it without that analysis simply moves the risk somewhere else in the process, often somewhere less visible.

Apple's public face verification research, along with NIST's broader testing program, both illustrate the same underlying pattern described throughout this article: published accuracy figures describe performance on curated, high-quality datasets, not on whatever image happens to come out of your camera on a given day. That distinction matters whether the comparison is running inside a consumer device or inside a forensic pipeline, the four-step logic of quality, score, threshold, and review applies either way, because the physics of image quality don't change based on who built the software.

Devices that capture the probe image are as much a part of the accuracy story as the matching software itself, since a device with a lower-resolution sensor, a fixed camera angle, or inconsistent lighting control will hand the algorithm a harder starting point regardless of how well-tuned that algorithm is. Evaluating a facial comparison result without accounting for the device that produced the original image is like grading a test without knowing which questions the student was actually asked.

Passcode-based verification and face-based verification are often deployed side by side precisely because facial comparison, however well-tuned, is a probabilistic tool rather than a certainty. Systems that fall back to a passcode or a secondary check when a facial comparison score is ambiguous are, in effect, building the human-review principle described earlier directly into the product rather than leaving it to a separate investigator. That fallback design is a practical acknowledgment that no confidence score, on its own, should be treated as the final word.

Taken together, these patterns point to the same conclusion this article has been building toward: a facial comparison result that seems to be working poorly is telling you something specific about quality, threshold, or review, not delivering a verdict on whether facial recognition technology in general can be trusted.

Face Matching: Why One Score Never Tells the Whole Story

Face matching is the general term for the whole pipeline described in this article, quality assessment, score generation, threshold comparison, and human review, not just the moment the algorithm returns a number. When someone says face matching failed, they usually mean the final output looked wrong, but the failure almost always traces back to one of the earlier stages in that chain. Understanding face matching as a process, rather than a single technical action, is what lets an investigator diagnose the actual point of failure instead of guessing.

Similarity Score: Reading the Number Correctly

A similarity score is the raw output of the comparison algorithm, and by itself it measures how alike two extracted templates are, nothing more. Treating the similarity score as a finished answer skips the quality and threshold context that determines whether that number means what it appears to mean. A careful reviewer always asks what conditions produced a given similarity score before deciding how much weight it deserves.

Face Match Results: What "Face Match" Actually Certifies

A face match label on a report means the algorithm's score cleared the threshold in place at the time, not that a human has confirmed identity beyond doubt. Calling something a face match without noting the threshold and the image quality behind it strips away exactly the context a reader needs to judge how much trust that label deserves. That's why careful documentation states the face match score alongside the quality conditions, rather than the label alone.

Percentage Figures: Why a Face Match Percentage Isn't a Probability

A face match percentage looks like a straightforward probability, but it's actually a similarity measurement shaped by both the underlying likeness and the quality of the two images being compared. When a report cites a percentage, it should also state the quality conditions that produced it, since the same two people can generate a very different percentage on a well-lit frontal photo than on a grainy angled frame. Reading a percentage without that context is the single most common way people over-trust a facial comparison result.

Photos and Their Role in Determining Match Quality

The photos fed into a comparison system set the ceiling on how reliable any resulting score can be, regardless of how sophisticated the underlying algorithm is. Two clear, frontal, well-lit photos give an algorithm its best chance at an accurate similarity measurement, while grainy or angled photos push the whole process toward the degraded conditions described earlier in this article. Anyone assessing a match result should ask what the source photos actually looked like before trusting the number attached to them.

Features Extraction: The Hidden Step Behind Every Score

Every comparison begins with features extraction, the process where the algorithm identifies measurable points on a face, spacing, contours, proportions, and converts them into a template for comparison. Poor image quality degrades the features extraction step before the comparison even happens, which is exactly why a low score can reflect a bad photo rather than a real difference between two people. Understanding features extraction as the true starting point of the pipeline helps explain why identical algorithms can produce very different scores on the same person.

Instant Face Match Score Tools: Convenience Versus Context

Consumer-facing tools that promise an instant face match score are built for speed, not for the quality-and-threshold analysis described throughout this article. An instant face match score can be a useful starting signal, but it should never substitute for the review process a serious case requires, because the tool behind it typically has no way to flag whether the input photos met any quality standard at all. Treating a quick score as casually informative, rather than as evidentiary, keeps expectations aligned with what these tools actually do.

Two Faces, One Comparison: What Closely Two Faces Align Actually Means

Every comparison ultimately reduces to two faces and a single number describing how closely two faces align across the measurable points the algorithm extracted. That number says nothing on its own about whether the two faces belong to the same person until it's read alongside the image quality and the threshold that generated it. Framing the task as two faces being measured against each other, rather than a yes-or-no identity check, keeps the limits of the technology in view.

Face Comparison Structure: How Facepair Analysis Is Organized

A facepair is simply the two images placed side by side for comparison, the probe and the reference, and the structure of that comparison matters as much as the algorithm running it. Good facepair analysis documents both images' quality conditions alongside the resulting score, so a later reviewer can reconstruct why the number came out the way it did. That structure is what turns a one-off comparison into something a case file can actually rely on.

A face matcher is only as trustworthy as the discipline wrapped around it: the quality checks run before comparison, the threshold chosen for the specific case, and the human review applied after the score appears. Teams that get your detailed similarity percentage from a vendor report should ask the same four questions this article has raised throughout, what was the image quality, what threshold produced this number, who reviewed it, and does the percentage match what a trained eye sees. That habit turns a single face comparison figure into a defensible finding rather than a number taken on faith.

Frequently asked questions

What resolution issue can secretly cut facial recognition accuracy in half?

When the inter-eye pixel distance in a probe image falls below 24 pixels, accuracy can drop by 50 percentage points before the algorithm even runs a fair comparison. This resolution level shows up constantly in real operational footage, and neither the algorithm nor the confidence score warns you when it happens, so the result looks identical to a perfect-quality match.

Does software for facial recognition give a final, conclusive answer on its own?

No. Software for facial recognition produces a similarity score, but that score is a signal, not a conclusion. In serious investigations, the algorithm's output is where the story starts, not where it ends, and it needs to survive four distinct human and technical filters before it belongs in a report, to a client, or in a courtroom.

Why can't a facial recognition match score be trusted without human review?

A match score can look identical whether the image quality was excellent or poor, since problems like low inter-eye pixel distance aren't flagged by the system. Human review is a step that cannot be shortcut, because it catches quality and threshold issues the confidence score alone does not reveal, protecting against costly, wrong conclusions.

Ready for forensic-grade facial comparison?

Full forensic reports with detailed similarity scoring. Results in seconds.

Run My First Search