How our numbers are made
Every statistic on this site is a claim with a definition, a computation and a denominator — and a credence mark: ● measured from records, ◐ inferred from the name graph, ○ machine-generated. This page spells each one out, including the ways they can be wrong. We publish our failure modes on purpose: you cannot trust a number whose limits you are not shown.
The build loop
Collect
Names are gathered from public sources at world scale, each with different structure and reliability.
- Official registries
- Scanned name books
- Phonebooks
- Corporate registries
- Web scraping
Recorded bearers & record counts
A "recorded bearer" is one person-name record in the Mondonomo corpus, aggregated from official registers, reference works and public datasets. Records are not people: the same person can appear in several sources.
- How it's computed:
- We count distinct records per (name form, country) after per-source deduplication, then sum across the sources that attest the form.
- Denominator:
- All records attesting the same name form in the same country.
Known failure modes
- Sources overlap: a person present in two datasets counts twice at corpus level, so raw record counts overstate people. The register-counted bearer figures do not have this problem; raw "recorded" counts do.
- Coverage is uneven — countries with strong official registers (e.g. Nordic statistics offices) look overrepresented next to countries where we rely on smaller public lists.
Counted bearers & "1 in N"
How many people in a country a national register counted under the name, and its reciprocal against the national population: "1 in N people".
- How it's computed:
- A national register published the count; we divide by the national population (CLDR figures) to get the share. We do not estimate it. Where no register we accept counts names in a country, we publish neither figure.
- Denominator:
- National population, not our record count.
Known failure modes
- It is a floor, not a total. Where a name is spread over several spellings we take the largest register entry rather than summing them, because two spellings of one name are largely the same people spelled twice.
- A register counts what it counts. Some are full population stocks, some are top-N lists, and the figure carries whatever the register left out — so it understates before it overstates.
- Mixing count types is a known onomastic trap — living-population (stock) counts must never be confused with per-year birth registrations (flow). Registers whose totals are cohort-scale are refused rather than read as populations.
Recorded gender
Whether the name is recorded as a male name, a female name, or for both. We state a direction only — we do not publish a percentage split.
- How it's computed:
- Bearer-weighted male/female counts for the given-name form (and across its cluster for cluster-level statements), reduced to a direction. A minority share under 25% is treated as record noise rather than usage.
- Denominator:
- Only records that carry a gender label — for most names the majority carry none.
Known failure modes
- The gender on a record belongs to the person, and it is attached to every given name on that record. Where a name string also carries a parent’s or spouse’s given name — a patronymic in Greek or Pakistani records, a Spanish compound such as María José — an unambiguously male name accumulates female records and vice versa. Measured across our sources this runs from about 6% (Germany) to 35% (Pakistan) of the gendered records, which is why we publish no percentage.
- The gendered subset is not a random sample of the corpus: it is dominated by the markets with the highest contamination, so the clean registry-style sources are the ones that drop out of the denominator.
- Gender conventions differ by culture and era; a name strongly male in one country can be female elsewhere. The statement shown is global unless a country is stated.
Related names (the name graph)
Names connected to this one in the name2name graph: same-name-in-another-script, cognates and etymological relatives, variants and short forms, feminine/masculine counterparts.
- How it's computed:
- Edges are inferred by combining etymological sources with statistical evidence (shared bearers, co-occurrence, transliteration models), then filtered by per-relation importance floors.
- Denominator:
- All candidate edges for the name; only edges above the importance floor are shown.
Known failure modes
- Statistical edges can connect merely-popular names: an earlier build linked Klaus to Michael with high confidence simply because both are extremely common. Importance floors and per-relation caps exist precisely because of such cases — some noise still gets through.
- Relations are marked ◐ inferred, not ● measured: an edge is a model claim, not a documented etymology, unless the etymology section cites sources.
Meaning & origin
A recorded meaning (gloss) or derivation of the name, with the sources that attest it.
- How it's computed:
- Aggregated from etymological reference sources (each with an authority tier); shown only when a source actually records it.
- Denominator:
- Only ~1–2% of name forms have a verified etymology in our sources — the honest fallback line on the other pages is deliberate.
Known failure modes
- Folk etymologies circulate widely; where sources disagree we prefer the higher-tier source and omit contested claims rather than average them.
- A gloss can be language-specific: the same spelling can be a Norse-derived name in one language and an unrelated word in another.
Scripts & transliterations
How the name is written in other scripts: ● measured rows are spellings attested in records; ○ machine rows are model transliterations.
- How it's computed:
- Measured rows come from cross-script links in the corpus; machine rows from our WFST transliteration router trained on attested name pairs.
- Denominator:
- Scripts with either an attested spelling or a model route for the source language.
Known failure modes
- Machine transliterations follow the statistically dominant convention; personal or regional romanization preferences (Đorđe → Djordje vs Dordje) may differ.
- For rare names the model generalizes from similar names and can produce plausible-but-unattested spellings — that is why the ○ mark never upgrades to ●.
Pronunciation (IPA)
A machine grapheme-to-phoneme transcription in the International Phonetic Alphabet.
- How it's computed:
- Generated by our g2p models for the page language; not a recording of a human speaker.
- Denominator:
- The page language only — the same spelling sounds different in other languages.
Known failure modes
- Names violate regular spelling rules more than ordinary words; the model can regularize an irregular family pronunciation.
- One transcription is shown even where variants exist (dialects, generational shifts).
Rank ("the 3rd most common surname in Croatia")
The name's position among all names of the same kind — given names ranked against given names, surnames against surnames — in a country or worldwide. 1 = most common.
- How it's computed:
- Names are ranked by recorded count. The ranked unit is the whole name group where one exists, so every spelling and script of the same name counts once, together; a name with no recorded variants ranks as itself. Names with fewer than 5 records in a country do not enter that country's ranking.
- Denominator:
- Every name of that kind recorded in the country (or worldwide) above the 5-record floor — published beside the rank, because "#3" on its own cannot be checked.
Known failure modes
- Rank inherits our coverage. Where our record base for a country is thin or skewed toward particular sources, the ordering of nearby ranks is unstable even when the top few are right.
- Rank is by records, not by living people: it does not correct for sources that overlap, nor for a name's age profile. A name common among the deceased outranks its share of the living population.
- Only clustered names carry a rank today. A name with no recorded spelling variants is ranked in the denominator but is not yet shown its own rank on its page.
- Ranks beyond 1,000 are not shown at all: at that depth the difference between neighbouring positions is smaller than the error in our counts.
Data snapshots & freshness
Every page renders from a versioned build of the name graph and statistics, not from live scraping.
- How it's computed:
- Offline builds fuse the corpus, the name2name graph and reference sources into versioned datasets; the site reports the build date it serves.
- Denominator:
- The snapshot date shown on the statistics page applies to all numbers site-wide.
Known failure modes
- Numbers change between builds as sources are added and deduplication improves — cite Mondonomo with the snapshot date.