The world's names, by the numbers

Generated 2026-08-29 from 715,907,523 deduplicated attestations. Every figure here has one definition, published on the methodology page.

How big is it? Four honest answers.

One corpus counted four ways. The numbers differ because the units differ — pick a rung to see what it means and what falls away to reach it.

Named-entity spellings

0

Unique spellings labelled person, organization or place.

  1. Deduplicate: a billion rows collapse into distinct (spelling, language) pairs.(−17.6B)

  2. Drop the language axis — the same spelling in twelve languages is one string.(−630.3M)

  3. Keep only spellings labelled person, organization or place; dictionary words and unknowns fall away.(−66.2M)

Bar lengths are logarithmic — the top rung is about 96× the bottom one.

191.2M distinct spellings in scope

A truly global graph

Named-entity spellings per country. Click a country to break it down; compare two side by side.

fewermore all names

Click any country to break it down.

Languages & writing systems

Across the whole corpus.

Top languages

  • English91.1M
  • Hindi37.6M
  • Marathi36.8M
  • Tamil36.4M
  • Urdu35M
  • Bangla34.8M
  • Malayalam34.5M
  • Telugu34.5M
  • Gujarati34.5M
  • Kannada34.5M
  • Chinese21.1M
  • French18.9M
  • Spanish13.2M
  • Russian12.8M
  • Japanese9M
Show as table
Distinct spellingsdistinct spellings
NameDistinct spellings
English91,108,264
Hindi37,644,381
Marathi36,843,180
Tamil36,404,267
Urdu35,046,062
Bangla34,800,900
Malayalam34,532,765
Telugu34,519,387
Gujarati34,505,380
Kannada34,504,171
Chinese21,123,024
French18,869,394
Spanish13,193,433
Russian12,842,447
Japanese8,985,977

Read these as “attested in a corpus tagged with this language”, not as disjoint sets. A name recorded in a multilingual country is tagged with each of that country's languages, which is why the Indic tags sit so close together — they largely share one inventory rather than holding separate ones.

Top scripts

  • Latin137.2M
  • Han23.3M
  • Cyrillic14.9M
  • Thai3.3M
  • Arabic3.2M
  • Hangul2.8M
  • Devanagari2.6M
  • Hebrew821.4K
  • Kana550.7K
  • Japanese444.3K
  • Greek436.6K
  • Georgian398K
Show as table
Distinct spellingsdistinct spellings
NameDistinct spellings
Latin137,200,796
Han23,316,077
Cyrillic14,906,490
Thai3,327,147
Arabic3,208,935
Hangul2,843,759
Devanagari2,554,944
Hebrew821,363
Kana550,716
Japanese444,265
Greek436,555
Georgian397,969

Corpus-wide — the filter above does not apply below

Where the names come from

Five source families, from web scrapes to official registries.

Scrapes & web sources
129.2M+
Corporate registries
0+
Official registries
1.5M+
Geographic registries
31.7M+
Reference & etymology
139.2M+

How many ways can one name be spelled in Latin?

Each non-Latin form takes several Latin spellings — an estimate, not a count.

  • Latin: 137.2M forms
  • CJK: 23.3M forms
  • Cyrillic: 14.9M forms
  • Thai: 3.3M forms
  • Arabic: 3.2M forms
  • Hangul: 2.8M forms
  • Devanagari: 2.6M forms
  • Hebrew: 821.4K forms
  • Kana: 550.7K forms
  • Jpan: 444.3K forms
  • Greek: 436.6K forms
  • Georgian: 398K forms
  • Ethi: 316.8K forms
  • Khmr: 228.5K forms
  • Armenian: 194.2K forms
  • Hira: 163.1K forms
  • Bengali: 152.1K forms
  • Mymr: 82.2K forms
  • Tamil: 66.4K forms
  • Mlym: 50.5K forms
  • Telu: 36.9K forms
  • Other: 32.5K forms
  • Gujr: 23K forms
  • Knda: 21.9K forms
  • Laoo: 21.4K forms
  • Sinh: 17.6K forms
  • Guru: 17.2K forms
  • Orya: 13.7K forms
  • Tibt: 6.3K forms
  • Mong: 2.7K forms
  • Bopo: 1.2K forms
  • Syrc: 1.2K forms
  • Tfng: 1.1K forms
  • Thaa: 1.1K forms
  • Cans: 976 forms
  • Cher: 628 forms
  • Tglg: 622 forms
  • Nkoo: 361 forms
  • Yiii: 355 forms
  • Goth: 194 forms
  • Copt: 175 forms
  • Olck: 125 forms
  • Mtei: 72 forms
  • Java: 49 forms
  • Sylo: 47 forms
  • Xpeo: 39 forms
  • Tavt: 34 forms
  • Brah: 28 forms
  • Batk: 24 forms
  • Adlm: 18 forms
  • Aghb: 18 forms
  • Bugi: 17 forms
  • Phnx: 16 forms
  • Kthi: 15 forms
  • Tale: 13 forms
  • Sund: 13 forms
  • Egyp: 13 forms
  • Lana: 12 forms
  • Talu: 12 forms
  • Bali: 9 forms
  • Runr: 9 forms
  • Phli: 8 forms
  • Ogam: 7 forms
  • Xsux: 7 forms
  • Lisu: 6 forms
  • Vaii: 6 forms
  • Newa: 5 forms
  • Ugar: 5 forms
  • Orkh: 4 forms
  • Cham: 4 forms
  • Merc: 3 forms
  • Bamu: 3 forms
  • Lepc: 3 forms
  • Linb: 3 forms
  • Mand: 2 forms
  • Tagb: 2 forms
  • Kali: 2 forms
  • Osge: 2 forms
  • Limb: 1 forms
  • Khar: 1 forms
  • Armi: 1 forms
  • Perm: 1 forms
  • Ital: 1 forms
  • Prti: 1 forms
  • Sogd: 1 forms
  • Glag: 1 forms
  • Wara: 1 forms
  • Avst: 1 forms
  • Sind: 1 forms
  • Tang: 1 forms

Anatomy of a name

Every form is classified into an entity and a name-part role.

Plus 75.8M ambiguous / dictionary tokens (not counted as named entities).

  • People: Given 67782016, Surname 56398042, Patronymic 4442663, Title 1228942, Particle 651595, Full name 594
  • Organizations: Org name 40191189, Org word 16558
  • Places: City 38885681, Region/Country 32716584
  • Other: Other 70966011, Common word 4688766, Profession 159165, Stopword 252

Names across languages

Etymological links connecting names and concepts worldwide.

106M

Etymology links

74.2M

Cross-language links

10M

Name roots

interlingual-equivalent-of
73.3M
named-after
18.8M
variant-of
8.5M
derived-from
2.9M
transliteration-of
859.5K
cognate-of
684.2K
compound-of
320.1K
feminine-of
137.9K
hypocorism-of
137K
pseudonym-of
133.8K
borrowed-from
130K
masculine-of
120.3K
opposite-gender-of
4.3K
contracted-from
869
nickname-of
4

How the graph is built

An iterative bootstrap: collect public data, train models, re-collect — repeat.

again & again

Collect

Names are gathered from public sources at world scale, each with different structure and reliability.

  • Official registries
  • Scanned name books
  • Phonebooks
  • Corporate registries
  • Web scraping