ganjoorandroid/tools
Anas Rashid c0cb6a8ab6 Add the Urdu dictionaries, and suggest near words when nothing matches
Two Urdu sources join the Persian ones. Wiktionary's Urdu extract gives Urdu
headwords glossed in English, and Urdu Wiktionary itself gives definitions
written in Urdu — the only source here that does. The latter is thin, about
3,100 usable entries out of 31,000 pages since many are stubs, but for a word
it carries an Urdu reader is better served by it than by a translation into
English: عشق comes back as شدید جذبۂ محبت، گہری چاہت، محبت، پریم، پیار.

Every definition now names the dictionary and its language pair, and lays out
in the direction its own script reads, so an Urdu definition is right-aligned
beside a left-aligned English one.

When nothing matches, the sheet offers near words ranked by how many letters
they share with what was looked up, drawn from an index range scan on the
leading letters rather than a scan of the whole table. خودکامی, which has no
entry, offers خودکامه — the lemma it wants.

Measured honestly: the Urdu sources add little coverage over Persian — nine
words from the English-glossed extract, two from Urdu Wiktionary, against
1,285 from thirteen poems. They are here because an Urdu reader wants Urdu,
not because they widen the net.

The suggestion ranking is verified against the built database rather than only
on device: خودکامی → خودکامه, شیرازی → شیراز, مشکلها → مشکل. On an API 36
emulator all four sources answer عشق with their labels.

Known rough edge: affix stripping across four languages can mislead. ناولها
reaches ناول, the Urdu for "novel", which is not what Hafez meant. The matched
headword is always shown, so it is visible rather than silent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 17:29:38 +02:00
..
build_dictionary.py Add the Urdu dictionaries, and suggest near words when nothing matches 2026-10-04 17:29:38 +02:00
README.md Add the Urdu dictionaries, and suggest near words when nothing matches 2026-10-04 17:29:38 +02:00

Building the dictionary

app/src/main/assets/dictionary.db.gz is generated, not hand-written. This rebuilds it:

pip install readmdict python-lzo          # lzo needs the C library: brew install lzo
curl -L -o fa.jsonl https://kaikki.org/dictionary/Persian/kaikki.org-dictionary-Persian.jsonl
curl -L -o daneshjoo.mdx \
  "https://raw.githubusercontent.com/0xdolan/Daneshjoo/main/Daneshjoo%20Dictionary/Daneshjoo%20Dictionary.mdx"
python3 build_dictionary.py
gzip -9 ganjoor-dictionary.db

Why two sources

Neither alone is enough. Measured against every distinct word in five real poems (Hafez ×2, Saadi's Golestan, Rumi's Masnavi, a Khayyam rubaʿi — 485 words):

coverage
Daneshjoo alone 71%
Wiktionary alone + its form index ~80%
Both, with the lookup chain 88%

13 of the 55 remaining misses are Arabic lines quoted inside Persian poems, so Persian coverage is about 91%.

Wiktionary's export is 93 MB, of which the definitions are 0.94 MB — the rest is inflection tables, etymology templates, IPA and descendants. The build keeps the definitions and the form→lemma index (149,589 pairs) and drops the rest. That index is what resolves conjugated verbs: افتاد → افتادن, بگشاید → گشودن, دانند → دانستن.

Urdu sources

curl -L -o ur.jsonl https://kaikki.org/dictionary/Urdu/kaikki.org-dictionary-Urdu.jsonl
curl -L -o urwikt.xml.bz2 \
  https://dumps.wikimedia.org/urwiktionary/latest/urwiktionary-latest-pages-articles.xml.bz2
bunzip2 -k urwikt.xml.bz2

The first gives Urdu headwords glossed in English. The second is Urdu Wiktionary itself, the only source here whose definitions are written in Urdu — thin (around 3,100 usable entries out of 31,000 pages, many being stubs), but for a word it does carry an Urdu reader is better served by it than by a translation into English.

Licences

Each row carries its source, so attribution stays accurate and either source can be dropped without rebuilding the other.

  • Wiktionary — CC BY-SA 3.0. The generated database is therefore also CC BY-SA 3.0.
  • Daneshjoo — the repository states MIT. Note the underlying lexicon is a published Iranian dictionary, so that relicensing is worth verifying before relying on it.