CaraComp
CaraComp
Forensic-Grade AI Face Recognition for:
Get Started7-day refund guarantee**
privacyBy Cara Candelario

GDPR Biometric Data Rules: Consent, Access, and Compliance Explained

Facial Comparison vs. Face Harvesting: Why GDPR Treats Them Differently
An investigator compares two case-file photos, illustrating how gdpr biometric data rules govern lawful facial comparison.

Here's something that surprises almost everyone the first time they hear it: GDPR does not ban facial comparison. It doesn't ban algorithms that analyze faces. It doesn't even ban processing biometric data in investigations, full stop. What it regulates, with considerable precision, is the architecture around that processing. The collection method. The retention logic. The access controls. The purpose at the moment of capture.

That distinction, between the tool and the system the tool lives inside, is one of the most misunderstood ideas in investigative technology today. And it's costing practitioners real operational confidence they don't need to sacrifice.

TL;DR

Under EU law, the legal risk in facial comparison isn't the algorithm, it's the collection architecture around it. Tightly scoped, case-bound comparison of photos you lawfully hold occupies fundamentally different legal territory than mass biometric harvesting.

The myth goes something like this: "Faces are biometric data. Biometric data is sensitive data. Sensitive data is restricted under GDPR. Therefore, any AI that processes faces is legally dangerous." Follow that chain of reasoning and you'd conclude that uploading two passport photos from a case file is roughly equivalent to scraping 30 million faces from social media. Which is, to put it gently, nonsense.


What the Fireflies.ai Lawsuit Actually Teaches Us

Earlier this year, a lawsuit against Fireflies.ai, an AI meeting assistant that records, transcribes, and analyzes calls, started generating serious attention from privacy lawyers. The core concern wasn't that it used AI. It was what the tool actually did with biometric data: captured voice and potentially facial data from multiple parties simultaneously, often without granular individual consent from everyone in the room, retained that data persistently, and processed it across unrelated sessions and users.

That's the architecture that draws legal fire. As Epstein Becker Green analyzed in their breakdown of the case, the problems compound when biometric data is captured broadly, stored indefinitely, and where third parties have no meaningful ability to opt out of collection. The issue isn't the comparison. It's the continuous, uninvited, multi-party harvest. This article is part of a series, start with Stress Test Facial Comparison Method Against Deepf.

Now contrast that with an investigator who holds two photographs from a lawfully obtained case file and runs a facial comparison to determine whether the same person appears in both images. No new database is created. No third-party faces are swept up. The processing is closed at the end of the task. That's not a smaller version of what Fireflies.ai was doing. It's a categorically different operation, the same way a doctor comparing two X-rays from the same patient's folder is categorically different from photographing everyone entering a hospital to build a predictive health-risk profile. Same imaging principle. Completely different legal and ethical universe.

"The key question is not whether biometric data is processed, but whether the processing is proportionate to the purpose, limited to what is necessary, and accompanied by appropriate safeguards, including access controls and retention limits." Analysis framework, Skadden, Arps, Slate, Meagher & Flom LLP, EU GDPR Decisions Commentary

GDPR's Biometric Access Rules: What Investigators Need

Article 9 of GDPR restricts "biometric data processed for the purpose of uniquely identifying a natural person." That clause is doing a lot of work. The phrase "for the purpose of uniquely identifying" is not decorative, it signals that the intent and scope of the processing determines its legal classification, not simply whether a face appears in a dataset.

Here's where it gets genuinely interesting. The European Data Protection Board has issued guidance clarifying the treatment of pseudonymised biometric data, situations where re-identification risk is controlled, access is restricted to authorized personnel, and the processing is bounded by a defined purpose. In those conditions, the data occupies different legal territory than freely accessible, broadly identifiable personal data. The risk calculus changes. The compliance posture changes.

Article 9
GDPR restricts biometric data processed for the purpose of uniquely identifying a person, the operative phrase is about intent and scope, not about whether faces appear in data at all
Source: General Data Protection Regulation, EU 2016/679

For investigators, this translates into four concrete variables that determine defensibility:

1. Lawful basis for holding the original photos. If you have a subject's image because they submitted it during an employment application, because it's part of a court-ordered disclosure, or because you obtained it through proper legal channels, you already hold it lawfully. That matters from the moment the image enters your possession, not just at the moment of comparison.

2. Purpose limitation. The comparison must serve a defined investigative purpose. "We might need this later" is not a purpose. "We are comparing these two images to establish whether Individual A in Document X is the same person as Individual B in Document Y, as part of Case Reference 2024-0471" is a purpose. The specificity is the compliance. Previously in this series: Biometric Privacy Crackdowns Coming For Investigat.

3. Data minimisation at the algorithm level. This is the part most people miss, and it's where tool selection actually matters. A facial comparison run on two images that produces a similarity score and then discards the intermediate biometric template is architecturally different from a system that retains a facial embedding in a persistent database. If your tool isn't creating a new, lasting biometric profile, just answering the specific comparison question, you're operating at the minimum necessary data footprint. That's not a legal technicality; it's the core of what defensible facial comparison in investigations is designed to do.

4. Retention and access controls. Who can see the results? How long are they stored? Is access logged? The EDPB's guidance on pseudonymised data makes clear that strong access restriction and defined deletion timelines are among the clearest signals that processing is proportionate. Document these. Not for a regulator's benefit, for your own operational clarity.

The Four Pillars of Defensible Facial Comparison

  • âš¡ Lawful origin of imagesYou must hold the photos legitimately before comparison begins; that lawfulness carries through to processing
  • 📌 Defined, specific purposeComparison tied to a named case and a stated question, not a speculative or open-ended scan
  • 🔒 No persistent biometric databaseProcessing that answers a comparison question without creating a lasting profile satisfies data minimisation at its core
  • 📋 Access logs and retention policyDocumented controls on who sees results and when data is deleted are among the clearest markers of proportionate processing

Trusted by Investigators Worldwide
Run Forensic-Grade Comparisons in Seconds
Detailed facial comparison reports. Results in seconds.
Get Started
7-day refund guarantee**

When Biometric Data Processing Shifts Under GDPR

A landmark EU court ruling clarified something that had been ambiguous for years: pseudonymised data, where the re-identification risk is genuinely controlled and not merely claimed, can fall outside the strictest tier of GDPR personal data protections. Skadden's analysis of the ruling notes that the court examined not just whether data could theoretically re-identify someone, but whether re-identification was reasonably likely given the actual access controls and context in place.

That's a meaningful shift. It moves the legal question from a binary "is this personal data?" to a contextual "does this processing, in this environment, with these controls, create a real identification risk?" For investigators with well-documented workflows, specific case files, restricted access, defined deletion, the answer is increasingly defensible.

Meanwhile, White & Case's review of the EU Digital Omnibus signals that upcoming adjustments to the GDPR framework are likely to reinforce, not loosen, this context-sensitive approach. Purpose, proportionality, and provenance of data are becoming more central to enforcement thinking, not less. Up next: Biometric Privacy Crackdowns Small Investigators.

Look, nobody's saying this is simple. Biometric data in investigations carries real obligations, and those obligations deserve serious professional attention. But "serious attention" and "blanket avoidance" are not the same thing. Serious attention means understanding what your tools actually do to data after the comparison runs. It means documenting your lawful basis before you click compare, not after a complaint arrives.

Key Takeaway

GDPR evaluates why you process, what you retain, who can access it, and how long you keep it, not whether a face appears in your data. A tightly scoped comparison on two lawfully held images, with no persistent database created and documented controls in place, is architecturally and legally distinct from mass biometric harvesting. The myth that all face AI is the same under EU law isn't just wrong, it's operationally expensive for investigators who abandon defensible tools out of misplaced caution.

At CaraComp, the design principle behind facial comparison has always been that the tool should answer the investigative question, and then stop. No persistent templates. No shadow profiles. No aggregation across cases. That's not a marketing position. It's an architectural choice with direct compliance consequences, and it's exactly the kind of distinction regulators are now asking organizations to demonstrate, not just assert.

So here's the question worth sitting with: when a regulatory challenge to facial comparison technology eventually lands on an organization's desk, and given current scrutiny, it will, the decisive factor won't be which algorithm they used. It will be whether they can show exactly what happened to the data the moment the comparison finished. Can you answer that question about your current workflow right now, without looking anything up?

If the answer is no, that's not a technology problem. That's a documentation problem. And documentation problems are the easiest kind to fix.

Lawful Basis Under GDPR Biometric Data Rules

A lawful basis is simply the legal reason you're allowed to hold and process someone's information in the first place. Under GDPR, biometric data used to uniquely identify a person needs a stronger justification than ordinary personal data, usually explicit consent, a legal obligation, or a substantial public interest ground like preventing or investigating fraud. Investigators should be able to name their lawful basis in a single sentence, tied to the specific case, before any comparison runs.

Without a documented lawful basis, even a narrow, well-intentioned facial comparison sits on shaky legal ground. Write it down before you start, not after someone asks. That single habit closes most of the gap between "we meant well" and "we can prove it."

Processing Biometric Data: What Actually Changes the Risk

Processing biometric data covers anything you do with it: collecting it, comparing it, storing it, or deleting it. GDPR doesn't treat all processing the same way, a one-time comparison that ends in a discarded result carries far less risk than processing biometric data continuously to build a running profile of someone over time. The distinction investigators should track is duration and reuse, not just the initial act of processing biometric data.

If your workflow processes biometric data once, produces an answer, and stops, you're operating in the lower-risk category the regulation is built around. If the same data gets reused across unrelated matters, the risk profile shifts even though the underlying images haven't changed.

Biometric Information and Special Categories

Biometric information becomes legally sensitive under GDPR specifically when it's used to uniquely identify someone, a face used just to illustrate a document is treated differently from a face used to confirm identity. This is why Article 9 places biometric data among the "special categories" of personal data, alongside health records and other especially sensitive information. Special category status means extra safeguards apply, not that processing is automatically forbidden.

Knowing whether your use of biometric information crosses into unique-identification territory is the first practical step in any compliance review. If it does, the four pillars covered earlier in this article, lawful origin, defined purpose, minimisation, and access controls, become mandatory, not optional.

Biometric Recognition Versus Simple Comparison

Biometric recognition systems are typically built to scan large populations and flag matches automatically, often without any human decision at the point of capture. That's meaningfully different from an investigator manually comparing two specific images already in a case file. Regulators tend to scrutinize automated biometric recognition deployed at scale far more closely than bounded, human-directed comparison work.

Understanding this distinction helps investigators explain, in plain terms, why their tool isn't equivalent to a surveillance system. Recognition implies ongoing detection across a population; comparison implies answering one specific question about two known images.

Special Category Data and Explicit Consent

Special category data, the GDPR term for the most sensitive kinds of personal data, including biometric data used for identification, generally requires a higher bar than routine personal data processing. Explicit consent is one lawful basis for handling special category data, but it isn't the only one available to investigators; legal obligations and substantial public interest grounds can also apply depending on the case. What matters is that the basis is documented and matches the actual purpose of the processing.

Treating special category data casually, reusing it, storing it indefinitely, or sharing it without access controls, is where most avoidable risk lives. Treating it as what it is, a narrow and purpose-bound exception to ordinary data protection rules, keeps investigative work defensible.

Data protection thinking, in practice, is less about avoiding biometric data altogether and more about being able to explain, clearly and specifically, why you have it, what you did with it, and when it was deleted. GDPR gives investigators room to work with biometric data lawfully, the framework rewards documented discipline, not avoidance. Personal data of any kind, biometric or otherwise, earns stronger protection under GDPR the more sensitive and identifying it becomes, and biometric data used for recognition sits near the top of that scale.

GDPR Compliance and Consent in Practice

GDPR compliance for biometric data is less about a single checkbox and more about a chain of decisions that all have to hold up together: lawful basis, purpose, minimisation, and access control. Consent is one piece of that chain, but it is not the only lawful basis available, and it is not always the right one, consent that can be withdrawn at any moment is a poor foundation for an investigation that depends on stable, ongoing access to data. Where consent is used, it needs to be specific, informed, and freely given, not buried in a general terms-of-service agreement. Investigators who rely on a legal obligation or substantial public interest ground instead of consent should still document that choice with the same rigor, because GDPR compliance is judged on the reasoning, not just the outcome.

A privacy impact assessment is the practical tool most compliance teams reach for before biometric data processing begins at any scale. Running a privacy impact assessment forces someone to write down, in advance, what data is collected, why, who can access it, and how long it will be kept, the same four variables that determine defensibility throughout this article. Skipping that step doesn't make the risk disappear; it just means the first time anyone writes down the answers is after something has already gone wrong.

Data subject rights sit alongside consent as a second pillar of GDPR compliance that investigators sometimes overlook. A data subject has the right to ask what data an organization holds about them, including biometric data, and organizations need a real answer ready, not a shrug. Building a simple internal log of what biometric data exists, where it lives, and who can access it makes responding to a data subject request straightforward instead of a scramble.

Access their biometric data is exactly the phrase regulators use when describing what a data subject is entitled to request, and investigators should treat that entitlement as a design constraint from day one. If someone can access their biometric data records within your organization simply by asking a clear question and getting a clear answer, your access controls are already doing their job. If that request would trigger a frantic search through old case files, the gap is not legal, it's organizational, and it is fixable with better records.

Biometric data is personal information first, and a special category of personal information second, which is why every safeguard described above builds on ordinary data protection principles rather than replacing them. Recognizing that biometric data is personal information keeps the analysis grounded: the same lawful basis, purpose limitation, and retention discipline that apply to any personal data apply here, just with a higher bar attached.

Biometric data is classified as a special category under Article 9 specifically because it can reveal something as persistent and identifying as a person's face or voice pattern, not because the underlying image is inherently dangerous. Understanding that biometric data is classified this way for a specific legal reason, unique identification, helps investigators see why a photo used to illustrate a report sits in different territory from the same photo run through a facial comparison to confirm identity.

Specific technical processing steps are where a lot of the real compliance work actually happens, quietly, away from any policy document. The specific technical processing that turns a photograph into a comparable biometric template is the moment sensitivity spikes, because that step is what makes unique identification possible. Documenting exactly when that specific technical processing occurs, and what happens to the output afterward, gives investigators a factual answer instead of a guess if anyone ever asks.

Personal data resulting from a facial comparison, a similarity score, a match confidence, a simple yes-or-no determination, deserves the same retention discipline as the original photographs. Personal data resulting from any biometric process should be logged with the same care: who generated it, when, for what case, and when it will be deleted. Treating the output of a comparison as an afterthought is a common gap, and it's one of the easiest to close.

Biometric data becomes riskier the longer it sits without a clear reason for being kept, which is why retention limits matter as much as access controls. Once biometric data becomes disconnected from its original case purpose, sitting in a folder nobody remembers creating, it stops being a documented, defensible asset and starts being unmanaged risk. A simple deletion schedule, tied to case closure, keeps that from happening.

Data protection, at its core, is about matching the strength of a safeguard to the sensitivity of the data and the risk of the processing. GDPR does not ask investigators to avoid biometric data; it asks them to be able to explain, plainly, why they have it, what specific technical processing was applied, what personal data resulted, and when it will be gone. That explanation, written down before the work begins, is what turns a defensible practice into a demonstrable one.

Frequently asked questions

Does GDPR ban the use of biometric data in investigations?

No. GDPR does not ban facial comparison, algorithms that analyze faces, or biometric data processing in investigations outright. What it regulates is the architecture around that processing, including collection method, retention logic, access controls, and the purpose at the moment of capture. Tightly scoped, case-bound comparison of lawfully held photos occupies different legal territory than mass biometric harvesting.

What does GDPR biometric data actually mean under Article 9?

Article 9 of GDPR restricts biometric data processed for the purpose of uniquely identifying a natural person. The operative phrase, for the purpose of uniquely identifying, signals that intent and scope determine legal classification, not simply whether a face appears in a dataset. This distinguishes narrow, purpose-bound comparisons from broad identification systems.

What makes facial comparison legally risky under GDPR?

The legal risk in facial comparison isn't the algorithm itself but the collection architecture around it. Problems compound when biometric data is captured broadly from multiple parties without consent, stored indefinitely, and processed across unrelated sessions, as seen in the Fireflies.ai lawsuit. Closed, case-specific comparisons with no persistent database and defined access controls face a different risk calculus.

Ready for forensic-grade facial comparison?

Full forensic reports with detailed similarity scoring. Results in seconds.

Run My First Search