CaraComp
CaraComp
Forensic-Grade AI Face Recognition for:
Get Started7-day refund guarantee**
Podcast

Facial Recognition Bias: 34% Error Rate for Some Faces

Facial Recognition Bias: 34% Error Rate for Some Faces

Facial Recognition Bias: 34% Error Rate for Some Faces

0:00-0:00

This episode is based on our article:

Read the full article →

Facial Recognition Bias: 34% Error Rate for Some Faces

Full Episode Transcript


There's a facial recognition system out there that's ninety-nine percent accurate. And for one group of people, it fails a third of the time. Both of those things are true at once, and that's not a bug in the math. That's how the math was designed to be reported.


Almost half of American adults are already in a

Almost half of American adults are already in a facial recognition network used by law enforcement. As of two thousand sixteen, that was more than a hundred and seventeen million people. Nobody asked them. Nobody signed anything. If you've ever had a driver's license photo taken, there's a real chance your face is in there. That's unsettling, and I'm not going to pretend otherwise. But the thing that protects you isn't panic, it's knowing exactly what these accuracy numbers do and don't mean. So how does a system get called ninety-nine percent accurate while failing so many people?

The answer is averaging. An overall accuracy score is one number stretched across everybody the system was tested on. And an average can hide almost anything underneath it.

Picture a medical test that catches ninety-five percent of a disease. Sounds excellent. Now split the results by patient group. In one group it catches nearly everything. In another, it misses a third of the cases. The headline number never changed, but who the test actually works for just changed completely. That's what a match score does when nobody breaks it down by demographic.


Trusted by Investigators Worldwide
Run Forensic-Grade Comparisons in Seconds
Detailed facial comparison reports. Results in seconds.
Get Started
7-day refund guarantee**

Researchers have measured that gap

And researchers have measured that gap. Across commercial facial recognition systems, error rates for darker-skinned women ran up to thirty-four percent higher than for lighter-skinned men. The worst performance clustered consistently around one group, women who are Black, and between eighteen and thirty years old. Those aren't rounding errors. At the scale these systems run, that's false matches by the thousands.

Now, why would a vendor miss something that big? Because you literally can't see it unless you test at enormous scale and sort the results by who's in the photo. The National Institute of Standards and Technology, N.I.S.T., ran a study on a hundred and eighty-nine algorithms from ninety-nine different developers. They used more than eighteen million images of roughly eight and a half million people. A company testing on ten thousand photos wouldn't catch the disparity at all. It's not that they hid it. Most of them never looked hard enough to find it.

Here's the piece that surprised me most. That confidence number, zero point nine five, ninety-five percent, whatever the software shows you, doesn't mean the same thing on every face. Researchers found East Asian faces needed a higher threshold than Caucasian faces just to reach the same error rate. So the setting an operator picks once, and applies to everyone, is quietly stricter for some people and looser for others. For an investigator, that means one threshold can't be fair across a whole population. For everyone else, it means a number that looks objective on a screen isn't measuring the same thing from person to person.


The Bottom Line

And these failures aren't only about matching. Systems from major tech companies have repeatedly classified Black women as male. Not anonymous test subjects, Michelle Obama. Serena Williams. Sojourner Truth. When race and gender classification break down at the same time, the errors stack.

A match score isn't a measure of truth. It's a measure of how well a system performed under the conditions it was tested in. If it was validated mostly on lighter-skinned male faces, then its reliability on a darker-skinned woman's face isn't low, it's unknown. And unknown is not the same as ninety-nine percent.

So the three sentences to carry with you. An overall accuracy number is an average, and averages hide who the system fails. Tested properly, some of these systems missed darker-skinned women up to a third more often than lighter-skinned men. Which means the right question was never "how accurate is it", it's "who was it tested on?" You don't need a computer science degree to ask that question. You just need to know it exists. The written version goes deeper, link's below.

Ready for forensic-grade facial comparison?

Full forensic reports with detailed similarity scoring. Results in seconds.

Run My First Search