Falskt ID > Artiklar > The Layer That Falls Out: Why Surname Corpora Systematically Miss the Last Fifty Years

Den här artikeln är ännu inte översatt till Svenska — du läser originalet på English. Finns även på:Deutsch, English, Українська

The Layer That Falls Out: Why Surname Corpora Systematically Miss the Last Fifty Years

In Norway's national surname register, Mohamed is carried by 3,885 people. That places it at rank 99 — ahead of a great many names most people would call unremarkably Norwegian. Our Norwegian surname corpus, built from lists of "Norwegian surnames," did not contain it. Nor Khan (3,505, rank 106), nor Hassan (3,495, rank 107), nor Tran (3,253, rank 120).

That is not an interesting fact on its own. Any corpus is missing something. What makes it worth writing down is that we checked the same thing in five more countries on 18 July 2026, and the same shape came back every time — and that when we looked closely at what else was missing, the story turned out to be more complicated, and more useful, than "the data forgot the immigrants."

This article is about a blind spot in how name datasets get built. A companion piece, Migration Layers in Birth Cohorts, covers what happens on the time axis — how two generations of the same community produce opposite curves in the same registry. This one is about the snapshot: what is simply absent from the list.

The measurement

For each locale we took the national registry file, ranked it by bearer count, and asked a mechanical question: of the registry's top N surnames, how many does our corpus not contain? No judgement, no classification, just set difference.

LocaleRegistry sourceRegistry entriesCorpus entriesMissing from registry top-1000
en_GBONS58,2421,517305
fr_FRINSEE218,9831,731233
en_USUS Census 2010162,2542,064193
no_NOSSB3,62099323 (of top-200)
sv_SESCB411,7982,4450

Sweden's zero is not an error and not luck — it is the same measurement run against a corpus that had already been through this exact repair. It is in the table as the control.

What is actually missing

Here the intuitive story starts to break, and it is worth breaking properly.

Take the US list of 193. Grouped by the linguistic tradition the surname comes from:

GroupCountExamples
Spanish-language50Bonilla, Cisneros, Guevara, Villalobos
South/East Asian and Arabic15Yu, Choi, Kaur, Vang, Phan, Ahmed
Anglo and other European128Brennan, Koch, Nolan, Kline, Nielsen

Two thirds of what was missing from the American corpus was Brennan and Koch. The migration layer is a real and substantial part of the gap — a quarter of it is Spanish-language alone — but it is not the whole gap, and a corpus that only patched the migration names would still have been missing 128 surnames that no one would describe as recent arrivals.

Norway shows the same proportions in miniature. Of the 23 registry top-200 surnames absent from the corpus, 9 are migration-origin — Mohamed, Khan, Hassan, Tran, Ibrahim, Mohammed, Ahmad, Abdi, Hussein — and 14 are Hanssen (5,434, rank 68), Eliassen (4,141, rank 94), Hoel, Rønningen, Samuelsen, Solbakken, Aase, Sand, , Syversen, Berger, Øien, Enger and — pointedly — Andersson, the Swedish spelling, sitting in Norway with 2,538 bearers.

So the honest version of the finding is narrower than the headline, and it is this: a corpus assembled from "names of country X" is shallow, and the shallowness is not evenly distributed. It misses ordinary names because it is short. It misses migration-layer names because nobody thinks to put them on a list of Norwegian surnames — and those two failures compound at exactly the ranks where they are most visible.

The clearest single case is Britain. Miah is the highest-ranked surname missing from the British corpus, at rank 254 with 26,765 bearers. Behind it, Uddin (13,620, rank 537) and Li (8,643, rank 898). Meanwhile Patel (137,088, rank 24), Begum (78,156, rank 65), Singh (69,574, rank 78) and Kaur (48,488, rank 121) were all present. The corpus was not blind to South Asian surnames in general — only to the ones just below the level of household familiarity.

The French case is not what it looks like

France produced 233 missing names from the registry top-1000, including a conspicuous Portuguese-language group: Dos Santos (17,661, rank 281), Gonçalves (16,448, rank 314), De Oliveira (8,253, rank 780), De Sousa (7,694, rank 843), Teixeira (7,608, rank 857).

The obvious reading is a missing Lusophone layer. The obvious reading is wrong, or at least incomplete, because these were present: Da Silva (28,730, rank 137), Pereira (21,402, rank 202), Rodrigues (17,767, rank 275), Fernandes (17,316, rank 293).

The Portuguese-language layer was in the corpus. Parts of it fell out. And when we counted how many of the 233 missing French surnames contain a space, the answer was 5Dos Santos, De Oliveira, De Sousa, Le Borgne, Le Gal. Two of those are Breton. This is at least partly a multi-word surname handling problem, not a migration problem: names with a particle and a space survive some pipeline steps badly, and they do so regardless of what language they come from. (Le Gal has been in Brittany for considerably longer than fifty years.) The mechanics of particles and sorting are their own subject, covered in Name Prefixes and Sorting (companion article, not yet published).

Gonçalves and Teixeira have no space and were still missing, so the migration-layer effect is real in France too. But anyone reading the five-name Portuguese list as pure evidence of a migration blind spot would be reading a tokenizer bug as a demographic fact. We nearly did.

The registries have the same blind spot

The second half of this, and the part that cannot be fixed by making a corpus deeper: national registries under-record recent migration too, each for its own structural reason, and each in a knowable direction.

INSEE (France) tabulates surnames by decade of birth for births registered in France, 1891–2000. People born abroad are not in the file at all — roughly 10.7% of France's population. A first-generation arrival is invisible; only their French-born children appear. Every migration-origin figure from INSEE is therefore a floor.

SSB (Norway) publishes only surnames with 200 or more bearers. We can confirm this from the file itself: the minimum count across all 3,620 entries is exactly 200. Whatever sits below that — and a recently-arrived surname is disproportionately likely to sit below it — does not exist as far as the published statistics are concerned. This is a deliberate privacy floor, discussed in Privacy Thresholds in Name Registries; the point here is only that it is not demographically neutral.

ONS (Britain) is the sharpest case, because the failure is one of vintage. The British surname data in general circulation reflects roughly 2002. Poland and Romania joined the EU in 2004 and 2007. The consequence, straight from the file:

SurnameONS countONS rank
Nowak9766,705
Kowalski7907,949
Wójcik38713,153
Popescu11823,518
Ionescu5629,886
Mazurabsent

Mazur is one of the twenty most common surnames in Poland. It does not appear anywhere in a 58,242-entry British surname file. Against a present-day British population that includes several hundred thousand people of Polish origin and a large Romanian community, Popescu at 118 bearers is not a measurement — it is a timestamp. The file is not wrong about 2002; it is being read as though it were about now.

This is the trap that matters. A corpus builder who diligently sources everything from official national statistics, and refuses to add anything not in the registry, will still produce a dataset that is a generation behind — because they inherited the registry's cut-off along with its authority. Deepening the corpus does not help. The names are not further down the list. They are not on the list.

What we changed

The repair is unglamorous: take the registry top-N, set-difference it against the corpus, and add back everything missing at its real weight. Sweden's zero in the first table is what that looks like when it is finished — the same check that returns 305 for Britain returns nothing for Sweden, on a corpus of 2,445 entries against a 411,798-entry registry.

For the registry-vintage problem there is no clean fix, and we do not claim one. Where a registry is known to predate a major migration flow, the corpus inherits that gap, and the honest response is to say so in the data notes rather than to invent counts. A surname we cannot weight from a source is a surname we do not add.

What this does and does not claim

Every number above is a count of surname records. None of it measures ethnicity, nationality, citizenship, language spoken, or anything about any individual — France and several other states here do not collect ethnicity data at all, and nothing in this article is derived from a source that does. The groupings in the US table are our classification of the linguistic tradition a surname comes from, made for the purpose of describing a dataset gap; they are approximate, they are contestable at the margins, and Costa could reasonably sit in two of them.

The finding is about datasets, not about people: lists of "names of country X" encode the moment the list-maker's mental model was formed, and that moment is usually a few decades stale. The registries encode the moment the registry was compiled. Neither is neutral, both are correctable in one direction, and the arithmetic that finds the gap takes about four lines.


Data as of 2026-07-18

All set-difference measurements were computed on 18 July 2026 against the registry files listed below, as part of an audit of weighted name corpora for 64 locales. Ranks were recomputed from raw counts; where a registry contains ties, rank assignment can differ by a few positions from the publisher's own numbering — our build recorded Miah/Uddin/Li at 257/542/906 against the 254/537/898 computed here, and the discrepancy is tie-handling, not disagreement about the data.

Sources:

  • United Kingdom — ONS, surname counts (58,242 entries, minimum count 5). Data vintage approximately 2002; see the note on post-2004 migration above. <ons.gov.uk/&gt;
  • France — INSEE, Fichier des noms de famille, surnames by decade of birth, births registered in France 1891–2000 (218,983 entries). Excludes people born abroad. <insee.fr/fr/statistiques/353663…;
  • United States — US Census Bureau, Frequently Occurring Surnames in the 2010 Census (162,254 entries). <census.gov/topics/population/gene…;
  • Norway — SSB, surname statistics, StatBank table 12891 (3,620 entries, publication floor of 200 bearers, confirmed as the file minimum). <ssb.no/statbank/table/12891&g…;
  • Sweden — SCB, surname statistics (411,798 entries). SCB discontinued name-statistics production in 2024. <scb.se/&gt;

On sensitivity. This article describes gaps in datasets. It makes no claim about individuals or communities, records no ethnicity, and takes no position on migration. Every figure is a count of surname records in a public statistical file.

← Artiklar