Identità falsa > Articoli > The Privacy Threshold Problem in Name Registries

Questo articolo non è ancora stato tradotto in Italiano: stai leggendo l'originale in English. Disponibile anche in:Deutsch, English, Українська

The Privacy Threshold Problem in Name Registries

You look up a surname in a national registry and it isn't there.

What have you learned? Depending on which country's registry you queried, one of three completely different things: that the surname has zero bearers, that it has some number of bearers the agency won't tell you, or nothing whatsoever. The empty cell looks identical in all three cases. Only the registry's disclosure rule distinguishes them, and that rule lives somewhere other than the data file.

This matters the moment you act on the absence. If absence means zero, deleting the entry is correct. If absence means "below the threshold," deleting it destroys real data. Same evidence, same action, opposite verdict — decided entirely by a policy footnote.

We hit this while auditing surname corpora for 63 locales on 17 July 2026 (64 today), and nearly transferred a correct decision from one country to another where it would have been a serious error. Here is the comparison table we wish had existed beforehand.

The comparison

RegistryThresholdAbsence meansCan you clean by it?
Taiwan, MOINone — publishes surnames with 1 bearer (643 of them)Truly zeroYes
Denmark, DST<3 suppressed, but zero is reported separatelyDepends on the messagePer name
South Korea, KOSTAT≥5 bearers; the rest pooled into 기타 ("other")UnknownNo
Spain, INE≥20 bearers (statistical confidentiality rule)UnknownNo
France, INSEE≥30 births 1891–2000; the rest pooled as AUTRES NOMSUnknownNo
UkraineNo registry exists at allNothingNo
LatviaNo per-surname registry existsNothingNo

Seven registries, four different meanings for the same blank.

Denmark: the registry that answers the question directly

Denmark deserves the longest section because it does something no other source here does — it distinguishes the two states explicitly, in words. Danmarks Statistik operates a name oracle: you ask about a specific surname, and you get one of two answers.

DST responseWhat it meansCorrect action
Der er færre end 3 personer med efternavnet X1–2 bearers exist; the count is hiddenKeep it, weight 1
Der er ingen personer med efternavnet XZero bearersDelete it

"There are fewer than 3 people with the surname X" and "there are no people with the surname X" are different sentences, and Denmark is willing to say both. That single design decision converts a privacy threshold from an obstacle into a per-name adjudicator.

It is why we kept Thurah and Vagn — both fall inside the privacy band, both are real Danish surnames borne by one or two people — and why we deleted Pedersgaard. The registry does not merely fail to list Pedersgaard; it affirmatively states that nobody has it. That makes it a phantom, assembled from a plausible model (Peders- + -gaard) rather than observed in a population.

The argument rests entirely on the existence of two messages. If DST answered færre end 3 for zero as well — as a naive reading of "suppress counts below 3" would imply — Pedersgaard would be indistinguishable from Thurah and we would have had to keep it. The threshold would have swallowed the evidence. The disclosure policy, not the threshold value, is what makes Danish data usable this way.

Taiwan and South Korea: mirror images

These two are the reason this article exists. Neighbours, both East Asian, both with concentrated surname distributions, both publishing detailed official statistics — and on opposite sides of this problem, with nothing in the data's appearance to tell you so.

Taiwan: no floor at all

Taiwan's Ministry of the Interior publishes 643 surnames with exactly one bearer and 406 with two — down to Japanese-origin names like 八谷 and 大久保 held by a single person among 23 million. There is no privacy floor whatsoever.

That property turns the registry into a falsification instrument. Absence means zero bearers, full stop. So when we found 33 surnames in our Taiwanese corpus that the registry does not contain — 亓官, 漆雕, 巫馬, 段干, 百里, 東郭, 南門, 子車, 梁丘, 左丘, 東門, 西門, 南宮, 長孫, 宇文, 公孫, 萬俟, 聞人, 鍾離, 仲孫, 拓跋 and others — deleting them was correct. They had been taken from the 百家姓, an 11th-century classical text, and carried into a dataset meant to describe living people. Nobody in Taiwan has them.

South Korea: a floor at five

KOSTAT's census publishes only surnames with five or more bearers. Everything rarer is pooled into 기타 — "other". So a Korean surname missing from the published table might have four bearers, or two, or none, and the table cannot tell you which.

The Taiwanese deletion logic is therefore invalid in Korea. Not "less reliable" — invalid. It reasons from an absence that carries no information. We came close to applying it anyway, on the entirely reasonable-sounding grounds that we had just done exactly this, successfully, in a neighbouring country with a very similar-looking dataset. The difference is not in the data. It is in a disclosure rule published elsewhere.

And even in Taiwan, appearance decides nothing

The sharpest lesson came from inside the country where cleaning was permitted. 澹臺 looks precisely like the 33 ghosts: an archaic compound surname straight out of the 百家姓 — exactly the profile we had just deleted 33 entries for matching. It went onto the deletion list on that basis.

It has 2 bearers. It is in the registry. It stayed.

So did 司徒 (478), 上官 (294), 端木 (186), 諸葛 (180), 皇甫 (115), 慕容 (1) and 呼延 (1) — all archaic-looking, all real. Meanwhile 歐陽 has 7,653 bearers and 東郭 has zero, and no amount of philological expertise separates them by inspection.

Only the counter distinguishes them. Never the appearance. A registry with no privacy floor doesn't license judgement by eye; it licenses lookups.

Where nothing exists: Ukraine and Latvia

The third state — absence means nothing — is the least comfortable to sit with.

Ukraine has no surname registry available for verification at all. Not a thresholded one; none. Scholarly work publishes a top-100 and stops. So cleaning a Ukrainian surname tail reduces to deleting entries by how they sound, and we measured that heuristic: applied to deliberately selected "most suspicious" candidates, "sounds strange" produced a 21% false-positive rate — real surnames borne by identifiable people (an opera singer, a former Commander-in-Chief of the Armed Forces, a family of composers) sitting in the same tail as the genuine fabrications. The context makes it inevitable: Ukraine has 707,685 unique surnames, averaging 64 bearers each. In a distribution shaped like that, rare and strange is the baseline, not a defect signal.

Latvia adds a variant worth flagging, because it looks like the opposite. Latvia does publish a searchable per-name oracle — live, current, exactly the shape of the Danish tool. Query a surname and you get a count.

Except it is a database of given names only. Query Ozoliņš, one of the most common surnames in the country, and you get zero rows. Query Ivanovs and you get three people who have Ivanovs as a first name. The oracle answers a different question, fluently and with a straight face — and a registry that looks like the tool you need is more dangerous than an obvious absence, because it returns a number and the number is meaningless for your purpose.

So Latvia joins Ukraine: no per-surname registry, absence means nothing, cleaning forbidden.

Service categories that pretend to be names

A threshold has to put the suppressed remainder somewhere, and where it goes is usually a row that looks exactly like a name.

RegistryThe categoryWhat it isWhere it sits
Taiwan, MOI其他"Other" — the pooled remainderRank #147, 5,174 count
South Korea기타"Other" — everything below 5 bearersIn the surname table
France, INSEEAUTRES NOMSEverything under 30 births, 1891–2000In the national file

其他 at rank #147 is the one to think about. It is not buried in a tail where you'd never look. It sits comfortably inside any mechanical "take the top 200" slice, in a plausible position, formatted identically to every real surname around it. Take a top-N cut without an explicit exclusion and "Other" becomes the 147th most common surname in Taiwan in your dataset.

Sweden supplies the strangest variant. Its given-name registry carries the entire alphabet as names: M (305 bearers), A (120), S (65), I (22), K (17), B (5), Z (3). These are not Swedish naming practice; they are records where only an initial survived. The proof is structural rather than aesthetic — all 26 letters are present, which is not how a name distribution behaves.

But note what that argument does not license. Md (1,361 bearers) is not a letter of the alphabet; it is the Bangladeshi passport abbreviation of Muhammad, and 1,361 residents of Sweden have registered it as their legal called-name. It stayed. Likewise Thi (3,717) — in a Vietnamese parsed corpus that string is a notorious artifact, but in the Swedish register it is a deliberate registration. The rule that catches M must not catch Md, and "looks like an artifact" cannot tell them apart.

Such artifacts propagate in parsed corpora rather than legal registers. In Vietnamese frequency data, Y sits at rank 55 with 0.092% — and is not a surname at all. It is a male honorific of the Ê Đê and Gia Rai peoples, roughly "Mr". The proof is internal to the source: if Y were a surname, those peoples' actual clan names would sit near it. Instead Niê is at #210 (0.005%) and at #247. Y is eighteen times "more frequent" than the clan it supposedly belongs to, because a parser took the first token of every Ê Đê and Gia Rai man's name and filed it under "surname". The same source has Thị at #149 — a female middle name — plus Ka at #123 and A at #129.

Legal registers are structurally immune to that class: the field is "the surname on the document," so it cannot hold an honorific. Parsed corpora are not. Knowing which kind of source you hold tells you which artifacts to expect.

Thresholds also hide things that aren't thresholds

Some registries exclude whole populations, and that exclusion is not disclosed as a threshold because it isn't one.

France's INSEE file excludes people born abroad — about 10.7% of the French population. That is not a privacy rule; it is a systematic slice of the reference population, and it acts almost entirely on one layer. Surnames like Cissé or Benali are counted only through those bearers' French-born children. The file also counts births, not living bearers — a second unit shift on top of the first. The direction of the error is known; the magnitude is not, and multiplying by a guessed coefficient would be fitting, not correction.

The registry itself shows the layer arriving, by birth cohort, per 100,000:

Surname1941–19501991–2000Change
TRAORE0.1035.76×358
DIALLO0.3735.59×96
NGUYEN3.7655.35×15
MARTIN359.65280.93−22%

Sweden's given-name register offers a second variety: it counts living bearers of all ages, with no age breakdown available at all. Roughly 20% of its mass is children aged 0–17, so names fashionable since the 2010s are overstated for any adult-only application and names whose bearers have mostly died are understated. Direction known, magnitude unknown, no honest correction available.

Neither of these is a privacy threshold. Both are the same failure mode: the published population is not the population you assumed, and nothing in the file says so.

What to do with this

  • Find the disclosure rule before you touch the data. It is in a methodology note, not the data file, and it determines what an empty cell means.
  • Establish which of three states you're in. Absence = zero (act on it), absence = suppressed (don't), no source at all (don't, and say so).
  • Prefer registries that report zero separately. Denmark's two-message design is worth more than a lower threshold would be. The value of a threshold is not its number; it is whether zero is distinguishable from suppression.
  • Never port a cleaning decision across borders. Taiwan and Korea look alike and are opposites. The decision belongs to the registry, not to the data's appearance.
  • Exclude service categories explicitly. 其他, 기타, AUTRES NOMS and stray initials look like names and survive any mechanical top-N cut.
  • Check the unit and the reference population, not just the threshold. Births ≠ bearers. Citizens ≠ residents. All ages ≠ adults. Positions ≠ people.
  • When nothing exists, "no data" is the finding. Cleaning a tail by how it sounds has a measured 21% false-positive rate. That is not a heuristic; it is data destruction with a plausible cover story.

Data as of 2026-07-17

All registry behaviour described here was observed on 17 July 2026 in the course of building weighted surname corpora for the 63 locales the corpus held at that date; en_IE was added on 18 July 2026, bringing the set to 64. Thresholds are as published by each agency at that date. The Danish two-message behaviour, the Taiwanese single-bearer entries, the Latvian given-name-only oracle and the Vietnamese honorific artifact were each verified against the source rather than inferred. The Ukrainian and Latvian rows are negative findings — repeated attempts to locate a per-name national surname source found none available.

Sources:

  • Taiwan, Ministry of the Interior — surname rankings and counts by age band, 30 June 2023 (OGDL 1.0): data.gov.tw/dataset/126774
  • Denmark, Danmarks Statistik — name statistics and per-name lookup: dst.dk/da/Statistik/emner/bor… · open dataset of Danish forenames and surnames: sprogteknologi.dk/dataset/fornavne-og-ef…
  • South Korea, KOSTAT — 2015 Population and Housing Census, surname statistics (인구주택총조사 성씨 통계): kostat.go.kr
  • Spain, INE — surnames with ≥20 bearers, census of 1 January 2025; methodology note on the confidentiality rule: ine.es/daco/daco42/nombyapel/…
  • France, INSEE — Fichier des noms de famille 1891–2000 (≥30 births nationally; remainder pooled as AUTRES NOMS): insee.fr/fr/statistiques/353663…
  • Sweden, SCB — surnames and tilltalsnamn given names, 31 December 2022: scb.se
  • Latvia, PMLP — given-name database (surnames not covered): personvardi.pmlp.gov.lv · Latvia, CSP — 2021 census: stat.gov.lv
  • Ukraine — no per-name registry available; scholarly top-100 lists only
  • Vietnam — hoten.org VNTH01 (n = 1,682,729) and SG01 (n = 241,000) frequency datasets, CC BY 4.0 as declared in the page text; not a state registry

← Articoli