When a facial recognition match is wrong: what your staff should do
Every matching system is sometimes wrong. That is not a reason to avoid the technology, and it is not something to discover on the shop floor. This is the written procedure to agree before you switch anything on, and the record to keep when it happens.
SESam Erpik · Co-founder & CTO··8 min read
If you are evaluating facial recognition for a shop, someone on your board will eventually ask the right question: what happens when it gets it wrong? It is the question that decides whether a deployment is defensible, and it is usually answered with a reassurance about accuracy rates rather than a procedure. Accuracy rates are not the answer. A procedure is.
Here is our position, stated plainly because it shapes everything below. A facial recognition match in a retail setting should never be an accusation, an ejection, or an automatic anything. It is a prompt that says to a person: this might be someone on your list, go and look. If a wrong match happens and the customer never knew, the system behaved correctly. That outcome is a design choice, and it is one you can specify.
Why wrong matches happen at all
A matching system compares a face captured on camera against the entries on the retailer's own list and returns a similarity score. A threshold turns that score into either "worth a look" or "nothing here". That threshold is the whole game.
Set it high and you get very few wrong matches, and you also miss people who are genuinely on your list. Set it low and you catch more of them, and you hand your team more faces that turn out to be strangers. There is no setting that gives you zero of both. Anyone who tells you their threshold eliminates false matches is describing a system that misses almost everybody, or is not being straight with you.
Threshold triage · your own watchlistIllustrative
Watchlist reference · #QE113Each detected face gets a similarity score against the closest entry on your list.
78%
Drag it. You are choosing what reaches a person, nothing more.
Detection 01similarity 61%Nothing raised
Detection 02similarity 72%Nothing raised
Detection 03similarity 84%Review required
Detection 04similarity 91%Review required
Below the threshold: nothing raised, nothing to review.
Review queue
2 of 4 sent for review
Detection 0384%
Detection 0491%
Awaiting a person
Confirm matchNot a match
That decision belongs to a person in your team, not to this page.
Illustrative scores. The slider changes what a human sees. It never changes what happens to anyone.
That trade-off is why the interesting question is never "how accurate is it". It is "what does your team do in the seconds after a face appears on screen, and what happens to a person who should never have been on that screen at all".
The two failures that look identical from the outside
It helps to separate them, because they need different corrections even though they produce the same bad afternoon for a customer.
The system surfaced the wrong face. The person in front of your colleague is not the person in the list entry. The fix is a threshold and review question, and the entry itself is fine.
The system was right about the face, but the entry should not have been there. Someone was added on thin evidence, or was added properly and should have come off months ago. The fix is a governance question, and no amount of model accuracy touches it.
The second failure is the one retailers underestimate. A perfectly accurate system that matches against a poorly maintained list will treat innocent people badly with total precision. If you only audit the technology and never audit the list, you have audited the easy half.
The procedure to write before you go live
This is the part to agree, in writing, with your DPO and your store teams before a single camera is enabled. It is short on purpose. A procedure nobody can remember on a busy Saturday is not a procedure.
Name who may act on a match. A specific role, not whoever is nearest the screen. Everyone else escalates. If the named person is not on shift, the answer is that nothing happens.
Define what acting means, and write down what it does not mean. It does not mean approaching someone and telling them the system flagged them. It does not mean asking them to leave. It does not mean following them round the shop. In most cases it means being aware and being present.
Agree the words. If a discreet approach is warranted, it is an ordinary customer-service greeting, the same one any customer would get. Nobody is ever told a computer identified them.
Make "it is not me" the end of it. If a customer says the system has the wrong person, your colleague accepts that on the spot. They do not argue, do not re-check the screen in front of them, and do not ask for identification. The cost of being wrong about a stranger is far higher than the cost of letting one genuine match go.
Log every wrong match, with the same fields every time. This is the record that proves your system is being governed rather than merely running.
Correct the list the same day. If the entry was wrong or is out of date, it comes off. Give one named person the authority to remove an entry without a committee.
Tell the person how to complain, and make it a real route. A named contact, a response time you actually meet, and a clear statement of their right to complain to the ICO if you do not resolve it.
Review your wrong matches on a fixed schedule. Monthly is reasonable. If nobody is reviewing them, you will not notice the store or the camera or the entry that is generating most of them.
What to record when it happens
Keep this deliberately boring and identical every time. The value is in being able to answer a regulator, a complainant or your own board with a file rather than a recollection.
Date, time, store and camera.
Who reviewed the match and in what role.
That it was a match prompt, and the outcome the reviewer recorded.
Whether the customer was approached at all, and if so what was actually said.
Which of the two failures it was: wrong face, or an entry that should not have been live.
What was done to the list afterwards, by whom, and when.
Whether the person was told how to complain.
On our own platform this is why the audit log is append-only. A record you can quietly edit after a complaint arrives is not evidence of good governance, it is the opposite, and any vendor should be able to tell you plainly whether theirs can be edited.
What the person on the other side is entitled to
Biometric data used to identify someone is special category data under UK data protection law, which is why this deployment carries obligations an ordinary CCTV install does not. In practice, a person who believes they have been wrongly flagged can ask what personal data you hold about them, ask you to correct it if it is wrong, ask you to erase it, object to the processing, and complain to the ICO.
There is a second point worth understanding, because it is the legal backbone of everything above. Following the Data (Use and Access) Act 2025, the UK GDPR provisions on automated decision-making turn on whether there was meaningful human involvement in the decision, and they place tighter restrictions where a significant decision rests on special category data. Biometric identification is special category data. A design in which a person genuinely reviews and decides, rather than rubber-stamping whatever the screen says, is not a nice-to-have in that framework. It is the thing that keeps you on the right side of it.
"Meaningful" is doing real work in that sentence. A colleague who confirms every prompt without looking is not meaningful human involvement, whatever your process document says. That is a training and workload question as much as a legal one, and it is a good reason not to run a threshold so low that your team learns to click through.
What to ask a vendor before you sign
Can the system take any action on a person without a human confirming first? The answer you want is no, and you want to be shown where that is enforced rather than told it is policy.
Who can remove an entry from the list, and how long does it take?
Is the audit log append-only, and can you show me an attempted edit being refused?
What is recorded when a reviewer dismisses a prompt, as opposed to confirming one?
Does my watchlist ever leave my own tenancy, or get shared with other retailers?
What is the default retention on an entry, and what happens to it when we stop using the product?
What does the reviewer actually see at the moment of decision, and what are they not shown?
The fifth one matters more than it looks. A list that belongs to you and never leaves your tenancy means a wrong entry is a mistake you can correct in an afternoon. A list shared across an industry scheme means the same mistake follows someone into shops that have never met them, and you cannot fix it alone. Both models exist in UK retail today. They are different products with different risks, and a buyer should know which one they are being sold.
The honest summary
Facial recognition in a shop is defensible when a wrong match is a non-event: a colleague looks, sees a stranger, dismisses it, and the customer finishes their shopping without ever knowing. It is indefensible when a wrong match becomes an accusation. The technology does not determine which of those you get. Your procedure does, and it needs to exist on paper before it is needed in a store.
What should a shop do if facial recognition identifies the wrong person?
Nothing should have happened to them yet, because a match is a prompt for a colleague to look rather than a decision. If a customer has been approached and says it is not them, staff should accept that immediately without arguing or asking for identification, log it as a wrong match, correct or remove the list entry the same day, and tell the person how to complain. A named person should have authority to remove an entry without waiting for a committee.
Can facial recognition get it wrong?
Yes, and any vendor who says otherwise is describing a system tuned so high that it misses almost everyone. Matching produces a similarity score, and a threshold turns that score into a prompt or silence. A higher threshold means fewer wrong matches and more missed people; a lower one means the reverse. There is no setting with zero of both, which is why the procedure for handling a wrong match matters more than the accuracy figure.
What rights do I have if a shop wrongly flags me on facial recognition?
Biometric data used to identify someone is special category data under UK data protection law. You can ask the retailer what personal data they hold about you, ask them to correct it if it is wrong, ask for it to be erased, object to the processing, and complain to the Information Commissioner's Office if the retailer does not resolve it. The retailer should give you a named contact and a response time.
Does a person have to review every facial recognition match?
In a responsibly designed retail deployment, yes. Following the Data (Use and Access) Act 2025, the UK GDPR provisions on automated decision-making turn on whether there was meaningful human involvement, with tighter restrictions where a significant decision rests on special category data such as biometrics. Review that is genuine rather than a rubber stamp is what keeps the deployment on the right side of that framework.
What is the difference between a wrong match and a wrong watchlist entry?
A wrong match means the system surfaced someone who is not the person in the entry, which is a threshold and review problem. A wrong entry means the match was correct but the person should never have been on the list, or should have come off it, which is a governance problem. Both look the same to the customer and both need to be logged, but only the first is fixed by better technology.
One email a month. The post-of-the-month, the retail-trends summary, and one customer-success snippet. No sales pitches, no event invites. Opt out in one click.