Facial Recognition Bias: 34% Error Rate for Some Faces
Facial Recognition Bias: 34% Error Rate for Some Faces
This episode is based on our article:
Read the full article →Facial Recognition Bias: 34% Error Rate for Some Faces
Full Episode Transcript
There's a facial recognition system out there that's ninety-nine percent accurate. And for one group of people, it fails a third of the time. Both of those things are true at once, and that's not a bug in the math. That's how the math was designed to be reported.
Almost half of American adults are already in a
Almost half of American adults are already in a facial recognition network used by law enforcement. As of two thousand sixteen, that was more than a hundred and seventeen million people. Nobody asked them. Nobody signed anything. If you've ever had a driver's license photo taken, there's a real chance your face is in there. That's unsettling, and I'm not going to pretend otherwise. But the thing that protects you isn't panic, it's knowing exactly what these accuracy numbers do and don't mean. So how does a system get called ninety-nine percent accurate while failing so many people?
The answer is averaging. An overall accuracy score is one number stretched across everybody the system was tested on. And an average can hide almost anything underneath it.
Picture a medical test that catches ninety-five percent of a disease. Sounds excellent. Now split the results by patient group. In one group it catches nearly everything. In another, it misses a third of the cases. The headline number never changed, but who the test actually works for just changed completely. That's what a match score does when nobody breaks it down by demographic.
Researchers have measured that gap
And researchers have measured that gap. Across commercial facial recognition systems, error rates for darker-skinned women ran up to thirty-four percent higher than for lighter-skinned men. The worst performance clustered consistently around one group, women who are Black, and between eighteen and thirty years old. Those aren't rounding errors. At the scale these systems run, that's false matches by the thousands.
Now, why would a vendor miss something that big? Because you literally can't see it unless you test at enormous scale and sort the results by who's in the photo. The National Institute of Standards and Technology, N.I.S.T., ran a study on a hundred and eighty-nine algorithms from ninety-nine different developers. They used more than eighteen million images of roughly eight and a half million people. A company testing on ten thousand photos wouldn't catch the disparity at all. It's not that they hid it. Most of them never looked hard enough to find it.
Here's the piece that surprised me most. That confidence number, zero point nine five, ninety-five percent, whatever the software shows you, doesn't mean the same thing on every face. Researchers found East Asian faces needed a higher threshold than Caucasian faces just to reach the same error rate. So the setting an operator picks once, and applies to everyone, is quietly stricter for some people and looser for others. For an investigator, that means one threshold can't be fair across a whole population. For everyone else, it means a number that looks objective on a screen isn't measuring the same thing from person to person.
The Bottom Line
And these failures aren't only about matching. Systems from major tech companies have repeatedly classified Black women as male. Not anonymous test subjects, Michelle Obama. Serena Williams. Sojourner Truth. When race and gender classification break down at the same time, the errors stack.
A match score isn't a measure of truth. It's a measure of how well a system performed under the conditions it was tested in. If it was validated mostly on lighter-skinned male faces, then its reliability on a darker-skinned woman's face isn't low, it's unknown. And unknown is not the same as ninety-nine percent.
So the three sentences to carry with you. An overall accuracy number is an average, and averages hide who the system fails. Tested properly, some of these systems missed darker-skinned women up to a third more often than lighter-skinned men. Which means the right question was never "how accurate is it", it's "who was it tested on?" You don't need a computer science degree to ask that question. You just need to know it exists. The written version goes deeper, link's below.
Ready for forensic-grade facial comparison?
Full forensic reports with detailed similarity scoring. Results in seconds.
Run My First SearchMore Episodes
Facial Recognition Privacy Concerns: MSG Fined $30,000
A lawyer buys a ticket to a concert. She walks through the doors, and before she's even handed her phone to the scanner, a camera has already picked her face out of the crowd and flagged her for removal. Not because she d
PodcastChina Facial Recognition Rules: UK Scanned 4M Faces First
Picture yourself walking through a shopping street on a Saturday afternoon. In under five hours, a single police camera in central London scanned fifty thousand faces. Yours could have been one of them — and you'd never have known. <break ti
PodcastUK Digital Identity: 275 Firms Face One New Rulebook
A ninety-nine percent confidence score sounds like near-certainty. But run that same system across a database of a million faces, and it can hand you thousands of wrong answers. The number didn't lie. It just never meant
