ganjoorandroid/tools/README.md
Anas Rashid c0cb6a8ab6 Add the Urdu dictionaries, and suggest near words when nothing matches
Two Urdu sources join the Persian ones. Wiktionary's Urdu extract gives Urdu
headwords glossed in English, and Urdu Wiktionary itself gives definitions
written in Urdu — the only source here that does. The latter is thin, about
3,100 usable entries out of 31,000 pages since many are stubs, but for a word
it carries an Urdu reader is better served by it than by a translation into
English: عشق comes back as شدید جذبۂ محبت، گہری چاہت، محبت، پریم، پیار.

Every definition now names the dictionary and its language pair, and lays out
in the direction its own script reads, so an Urdu definition is right-aligned
beside a left-aligned English one.

When nothing matches, the sheet offers near words ranked by how many letters
they share with what was looked up, drawn from an index range scan on the
leading letters rather than a scan of the whole table. خودکامی, which has no
entry, offers خودکامه — the lemma it wants.

Measured honestly: the Urdu sources add little coverage over Persian — nine
words from the English-glossed extract, two from Urdu Wiktionary, against
1,285 from thirteen poems. They are here because an Urdu reader wants Urdu,
not because they widen the net.

The suggestion ranking is verified against the built database rather than only
on device: خودکامی → خودکامه, شیرازی → شیراز, مشکلها → مشکل. On an API 36
emulator all four sources answer عشق with their labels.

Known rough edge: affix stripping across four languages can mislead. ناولها
reaches ناول, the Urdu for "novel", which is not what Hafez meant. The matched
headword is always shown, so it is visible rather than silent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 17:29:38 +02:00

55 lines
2.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Building the dictionary
`app/src/main/assets/dictionary.db.gz` is generated, not hand-written. This rebuilds it:
```sh
pip install readmdict python-lzo # lzo needs the C library: brew install lzo
curl -L -o fa.jsonl https://kaikki.org/dictionary/Persian/kaikki.org-dictionary-Persian.jsonl
curl -L -o daneshjoo.mdx \
"https://raw.githubusercontent.com/0xdolan/Daneshjoo/main/Daneshjoo%20Dictionary/Daneshjoo%20Dictionary.mdx"
python3 build_dictionary.py
gzip -9 ganjoor-dictionary.db
```
## Why two sources
Neither alone is enough. Measured against every distinct word in five real poems (Hafez ×2,
Saadi's Golestan, Rumi's Masnavi, a Khayyam rubaʿi — 485 words):
| | coverage |
|---|---|
| Daneshjoo alone | 71% |
| Wiktionary alone + its form index | ~80% |
| **Both, with the lookup chain** | **88%** |
13 of the 55 remaining misses are Arabic lines quoted inside Persian poems, so Persian coverage
is about 91%.
Wiktionary's export is 93 MB, of which the definitions are 0.94 MB — the rest is inflection
tables, etymology templates, IPA and descendants. The build keeps the definitions and the
form→lemma index (149,589 pairs) and drops the rest. That index is what resolves conjugated
verbs: `افتاد → افتادن`, `بگشاید → گشودن`, `دانند → دانستن`.
## Urdu sources
```sh
curl -L -o ur.jsonl https://kaikki.org/dictionary/Urdu/kaikki.org-dictionary-Urdu.jsonl
curl -L -o urwikt.xml.bz2 \
https://dumps.wikimedia.org/urwiktionary/latest/urwiktionary-latest-pages-articles.xml.bz2
bunzip2 -k urwikt.xml.bz2
```
The first gives Urdu headwords glossed in English. The second is Urdu Wiktionary itself, the only
source here whose definitions are written **in Urdu** — thin (around 3,100 usable entries out of
31,000 pages, many being stubs), but for a word it does carry an Urdu reader is better served by
it than by a translation into English.
## Licences
Each row carries its `source`, so attribution stays accurate and either source can be dropped
without rebuilding the other.
- **Wiktionary** — CC BY-SA 3.0. The generated database is therefore also CC BY-SA 3.0.
- **Daneshjoo** — the repository states MIT. Note the underlying lexicon is a published Iranian
dictionary, so that relicensing is worth verifying before relying on it.