Technology

Data at a Glance

Each figure with the denominator it is counted over — the counts differ because the units do, not because they disagree.

191.2M+
Named-entity spellings

Unique spellings labelled person, organization or place.

257.4M+
Distinct spellings

Unique spellings across every label, dictionary words and unknowns included.

18.4B+
Records analyzed

Rows ingested from every source. Records are not people — one person can appear in many. method

263
Countries

ISO country codes with at least one attested form; 'unknown' excluded.

718
Official registries

Distinct source ids in the official-registry table.

1,032+
Languages with names

Language tags carrying at least one named entity.

106M+
Etymology links

Typed edges in the etymology graph, all predicates.

280.7M+
Possible romanizations

A projection, not a count: distinct non-Latin forms × a per-script romanization factor.

See the full picture → live statistics

Models

PNEUMA

Six-head name-understanding model (CNN-Deep v10) — predicts entity type, language, country, gender, name-part BIO tags, and the name's origin (decoupled from the script/language it is written in) across 140+ languages and 240+ countries.

Published

PNEUMA-DD

Token classification model (parse_dict) — classifies individual name tokens by role and computes per-token popularity and geographic distribution.

Published

MondoPhon

Phonological transliteration model — generates name variants across scripts and languages using learned phonological mappings.

In Development

AyutthayaAlpha

Thai-Latin script transliteration transformer — trained on 2.7M Thai-Latin pairs with CER 0.0047, covering Royal Thai General System of Transcription and phonetic variants.

Published

WFTS

Weighted finite-state transducer for transliteration — deterministic script conversion with probabilistic variant ranking.

Experimental

MondoGraph

Name data graph — interactive visualization of etymology, popularity, and transliteration relationships for a name token.

In Development

Technical Papers

Peer-reviewed and preprint publications describing the models and methods behind the Mondonomo platform. Click Quick preview to read in-page, or open on arXiv.

Efficient Multilingual Name Type Classification Using Convolutional Networks

Davor Lauc · Jan 2026

92.1% accuracy across 104 languages · 46× faster than fine-tuned XLM-RoBERTa · the published precursor to PNEUMA

arXiv

AyutthayaAlpha — A Thai-Latin Script Transliteration Transformer

Davor Lauc, Attapol Rutherford (Chulalongkorn University), Weerin Wongwarawipatr · Dec 2024

CER 0.0047 on RTGS romanization · trained on 2.7 million Thai-Latin name pairs

arXiv

PolyIPA — Multilingual Phoneme-to-Grapheme Conversion Model

Davor Lauc · Dec 2024

Mean CER 0.055 across 21 languages · the foundation for MondoPhon phonetic transliteration

arXiv