How our numbers are made

Every statistic on this site is a claim with a definition, a computation and a denominator — and a credence mark: measured from records, inferred from the name graph, machine-generated. This page spells each one out, including the ways they can be wrong. We publish our failure modes on purpose: you cannot trust a number whose limits you are not shown.

The build loop

again & again

Collect

Names are gathered from public sources at world scale, each with different structure and reliability.

  • Official registries
  • Scanned name books
  • Phonebooks
  • Corporate registries
  • Web scraping

Recorded bearers & record counts

A "recorded bearer" is one person-name record in the Mondonomo corpus, aggregated from official registers, reference works and public datasets. Records are not people: the same person can appear in several sources.

How it's computed:
We count distinct records per (name form, country) after per-source deduplication, then sum across the sources that attest the form.
Denominator:
All records attesting the same name form in the same country.

Known failure modes

  • Sources overlap: a person present in two datasets counts twice at corpus level, so raw record counts overstate people. The register-counted bearer figures do not have this problem; raw "recorded" counts do.
  • Coverage is uneven — countries with strong official registers (e.g. Nordic statistics offices) look overrepresented next to countries where we rely on smaller public lists.

Counted bearers & "1 in N"

How many people in a country a national register counted under the name, and its reciprocal against the national population: "1 in N people".

How it's computed:
A national register published the count; we divide by the national population (CLDR figures) to get the share. We do not estimate it. Where no register we accept counts names in a country, we publish neither figure.
Denominator:
National population, not our record count.

Known failure modes

  • It is a floor, not a total. Where a name is spread over several spellings we take the largest register entry rather than summing them, because two spellings of one name are largely the same people spelled twice.
  • A register counts what it counts. Some are full population stocks, some are top-N lists, and the figure carries whatever the register left out — so it understates before it overstates.
  • Mixing count types is a known onomastic trap — living-population (stock) counts must never be confused with per-year birth registrations (flow). Registers whose totals are cohort-scale are refused rather than read as populations.

Recorded gender

Whether the name is recorded as a male name, a female name, or for both. We state a direction only — we do not publish a percentage split.

How it's computed:
Bearer-weighted male/female counts for the given-name form (and across its cluster for cluster-level statements), reduced to a direction. A minority share under 25% is treated as record noise rather than usage.
Denominator:
Only records that carry a gender label — for most names the majority carry none.

Known failure modes

  • The gender on a record belongs to the person, and it is attached to every given name on that record. Where a name string also carries a parent’s or spouse’s given name — a patronymic in Greek or Pakistani records, a Spanish compound such as María José — an unambiguously male name accumulates female records and vice versa. Measured across our sources this runs from about 6% (Germany) to 35% (Pakistan) of the gendered records, which is why we publish no percentage.
  • The gendered subset is not a random sample of the corpus: it is dominated by the markets with the highest contamination, so the clean registry-style sources are the ones that drop out of the denominator.
  • Gender conventions differ by culture and era; a name strongly male in one country can be female elsewhere. The statement shown is global unless a country is stated.

Related names (the name graph)

Names connected to this one in the name2name graph: same-name-in-another-script, cognates and etymological relatives, variants and short forms, feminine/masculine counterparts.

How it's computed:
Edges are inferred by combining etymological sources with statistical evidence (shared bearers, co-occurrence, transliteration models), then filtered by per-relation importance floors.
Denominator:
All candidate edges for the name; only edges above the importance floor are shown.

Known failure modes

  • Statistical edges can connect merely-popular names: an earlier build linked Klaus to Michael with high confidence simply because both are extremely common. Importance floors and per-relation caps exist precisely because of such cases — some noise still gets through.
  • Relations are marked ◐ inferred, not ● measured: an edge is a model claim, not a documented etymology, unless the etymology section cites sources.

Meaning & origin

A recorded meaning (gloss) or derivation of the name, with the sources that attest it.

How it's computed:
Aggregated from etymological reference sources (each with an authority tier); shown only when a source actually records it.
Denominator:
Only ~1–2% of name forms have a verified etymology in our sources — the honest fallback line on the other pages is deliberate.

Known failure modes

  • Folk etymologies circulate widely; where sources disagree we prefer the higher-tier source and omit contested claims rather than average them.
  • A gloss can be language-specific: the same spelling can be a Norse-derived name in one language and an unrelated word in another.

Scripts & transliterations

How the name is written in other scripts: ● measured rows are spellings attested in records; ○ machine rows are model transliterations.

How it's computed:
Measured rows come from cross-script links in the corpus; machine rows from our WFST transliteration router trained on attested name pairs.
Denominator:
Scripts with either an attested spelling or a model route for the source language.

Known failure modes

  • Machine transliterations follow the statistically dominant convention; personal or regional romanization preferences (Đorđe → Djordje vs Dordje) may differ.
  • For rare names the model generalizes from similar names and can produce plausible-but-unattested spellings — that is why the ○ mark never upgrades to ●.

Pronunciation (IPA)

A machine grapheme-to-phoneme transcription in the International Phonetic Alphabet.

How it's computed:
Generated by our g2p models for the page language; not a recording of a human speaker.
Denominator:
The page language only — the same spelling sounds different in other languages.

Known failure modes

  • Names violate regular spelling rules more than ordinary words; the model can regularize an irregular family pronunciation.
  • One transcription is shown even where variants exist (dialects, generational shifts).

Rank ("the 3rd most common surname in Croatia")

The name's position among all names of the same kind — given names ranked against given names, surnames against surnames — in a country or worldwide. 1 = most common.

How it's computed:
Names are ranked by recorded count. The ranked unit is the whole name group where one exists, so every spelling and script of the same name counts once, together; a name with no recorded variants ranks as itself. Names with fewer than 5 records in a country do not enter that country's ranking.
Denominator:
Every name of that kind recorded in the country (or worldwide) above the 5-record floor — published beside the rank, because "#3" on its own cannot be checked.

Known failure modes

  • Rank inherits our coverage. Where our record base for a country is thin or skewed toward particular sources, the ordering of nearby ranks is unstable even when the top few are right.
  • Rank is by records, not by living people: it does not correct for sources that overlap, nor for a name's age profile. A name common among the deceased outranks its share of the living population.
  • Only clustered names carry a rank today. A name with no recorded spelling variants is ranked in the denominator but is not yet shown its own rank on its page.
  • Ranks beyond 1,000 are not shown at all: at that depth the difference between neighbouring positions is smaller than the error in our counts.

Data snapshots & freshness

Every page renders from a versioned build of the name graph and statistics, not from live scraping.

How it's computed:
Offline builds fuse the corpus, the name2name graph and reference sources into versioned datasets; the site reports the build date it serves.
Denominator:
The snapshot date shown on the statistics page applies to all numbers site-wide.

Known failure modes

  • Numbers change between builds as sources are added and deduplication improves — cite Mondonomo with the snapshot date.