Technology
Data at a Glance
Each figure with the denominator it is counted over — the counts differ because the units do, not because they disagree.
- 191.2M+
- Named-entity spellings
- 257.4M+
- Distinct spellings
- 18.4B+
- Records analyzed
- 263
- Countries
- 718
- Official registries
- 1,032+
- Languages with names
- 106M+
- Etymology links
- 280.7M+
- Possible romanizations
Unique spellings labelled person, organization or place.
Unique spellings across every label, dictionary words and unknowns included.
Rows ingested from every source. Records are not people — one person can appear in many. method
ISO country codes with at least one attested form; 'unknown' excluded.
Distinct source ids in the official-registry table.
Language tags carrying at least one named entity.
Typed edges in the etymology graph, all predicates.
A projection, not a count: distinct non-Latin forms × a per-script romanization factor.
Models
PNEUMA
Six-head name-understanding model (CNN-Deep v10) — predicts entity type, language, country, gender, name-part BIO tags, and the name's origin (decoupled from the script/language it is written in) across 140+ languages and 240+ countries.
PNEUMA-DD
Token classification model (parse_dict) — classifies individual name tokens by role and computes per-token popularity and geographic distribution.
MondoPhon
Phonological transliteration model — generates name variants across scripts and languages using learned phonological mappings.
AyutthayaAlpha
Thai-Latin script transliteration transformer — trained on 2.7M Thai-Latin pairs with CER 0.0047, covering Royal Thai General System of Transcription and phonetic variants.
WFTS
Weighted finite-state transducer for transliteration — deterministic script conversion with probabilistic variant ranking.
MondoGraph
Name data graph — interactive visualization of etymology, popularity, and transliteration relationships for a name token.
Technical Papers
Peer-reviewed and preprint publications describing the models and methods behind the Mondonomo platform. Click Quick preview to read in-page, or open on arXiv.
Efficient Multilingual Name Type Classification Using Convolutional Networks
Davor Lauc · Jan 2026
92.1% accuracy across 104 languages · 46× faster than fine-tuned XLM-RoBERTa · the published precursor to PNEUMA
AyutthayaAlpha — A Thai-Latin Script Transliteration Transformer
Davor Lauc, Attapol Rutherford (Chulalongkorn University), Weerin Wongwarawipatr · Dec 2024
CER 0.0047 on RTGS romanization · trained on 2.7 million Thai-Latin name pairs
PolyIPA — Multilingual Phoneme-to-Grapheme Conversion Model
Davor Lauc · Dec 2024
Mean CER 0.055 across 21 languages · the foundation for MondoPhon phonetic transliteration