# 0009 — Duplicate patients are prevented at registration, in two tiers

**Status:** Accepted · 2026-08-02

## Context

Duplicate patient records are the most expensive routine data problem a clinic
has. The allergy is on one record, the current medication on another, and the
clinician looking at the third has no way to know either exists. Catching it at
registration costs a receptionist a few seconds. Catching it later means merging
clinical data, which is genuinely hard to do safely and which no clinic wants to
sign off on.

The check has to survive how people actually type. "Mohammed" and "Mohamed",
"MOHAMMED", "Müller" and "Muller", "0771 234 567" and "+964 771 234 567" and
"00964-771-234-567" are all the same person and none of them match as stored.

The national identifier is the one field that reliably identifies a person — and
it is also the field we encrypt, which makes `WHERE national_id = ?` impossible.

## Decision

**Normalise before matching.** `PatientIdentity` derives `name_normalized` and
`phone_normalized` columns from the display values: case-folded, diacritics
stripped, Arabic alef and taa-marbuta variants unified, punctuation removed;
phone numbers reduced to digits with an international prefix or trunk zero
removed and matched on the trailing significant digits. The service is the only
writer, so the searchable copies cannot drift from what is on screen.

**Search the encrypted identifier through a blind index.** The national ID is
stored encrypted and accompanied by an HMAC of its normalised value. Exact-match
lookup hits an index on the HMAC; the plaintext is never stored and the HMAC
cannot be reversed without the application key. The ciphertext and the index are
written together or not at all — a record with one and not the other is
invisible to the check while looking complete on screen.

**Two tiers, because they warrant different answers.**

| Tier | Trigger | Behaviour |
|---|---|---|
| Conclusive | National identifier matches | Refused. Not overridable. The existing MRN and name come back with the error so the interface can offer to open it |
| Probable | Phone matches, or name *and* date of birth match | Paused. The candidates are shown and a person confirms this is someone else |

Name alone is never a trigger. In some catchments a third of the register shares
a surname, and an alert that fires constantly is an alert nobody reads.

## Consequences

**Good.** The realistic duplicate — same person, different spelling, different
phone — is caught by the identifier rather than by the name. Reception gets an
actionable message naming the existing record instead of a bare "duplicate".

**Bad.** Blind indexing supports equality only. Fuzzy or partial search over the
encrypted identifier is not possible and is not pretended to be.

**Bad.** Phone normalisation is not true E.164 — that needs a country context we
do not reliably have at the point of entry, and a library we chose not to ship.
Matching on trailing digits can in principle collide across country codes.
Within one clinic's catchment that is acceptable, and it is documented in the
code as the deliberate limit rather than dressed up.

**Not built.** Merging two records that are already duplicates. It is a real
requirement and a substantial one — it has to reconcile encounters, prescriptions
and invoices — and it belongs after the modules that own those records exist.
