ganjoorandroid/tools/README.md
Anas Rashid c548e6ec07 Add a tap-a-word dictionary, layering Wiktionary and Daneshjoo
Tapping a word in a poem now opens its meaning. Neither source covers enough
alone — Daneshjoo reaches 71% of real poem vocabulary — so both ship, each row
carrying its source so the credit stays attached and either can be dropped
later. Together with the lookup chain that reaches 88%, and 13 of the 55
remaining misses are Arabic lines quoted inside Persian poems.

Lookup widens until something matches: the word as written, the lemma it
inflects from, the word with an affix stripped, then the parts of a ZWNJ
compound. The lemma step is what makes classical verse readable, and it comes
from Wiktionary's 149,589 form->lemma pairs: افتاد is only findable as افتادن.
The Arabic definite article is stripped too, since poems quote Arabic.

Two things the build taught me. Wiktionary's 93 MB export is 0.94 MB of
definitions wrapped in inflection tables, etymology templates and IPA, so
tools/build_dictionary.py keeps the definitions and the form index and drops
the rest. And the packager silently gunzips .gz assets and strips the
extension, which shipped the database under a name the code wasn't opening —
every lookup failed silently until the APK listing gave it away.

Hit testing needed care as well: nastaliq is set with 2.4x leading, so most of
a line box is empty space, and a tap there clamps to the line's first
character — which made every tap return the opening word. Taps are now checked
against the baseline band, with the multipliers left as a knob for swapped
fonts. Definitions are English, so they read left-to-right inside the
otherwise right-to-left sheet.

Verified on the release build, where R8 could have broken the SQLite path:
tapping عشق returns both sources, and دانند resolves to دانستن.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 02:02:00 +02:00

41 lines
1.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Building the dictionary
`app/src/main/assets/dictionary.db.gz` is generated, not hand-written. This rebuilds it:
```sh
pip install readmdict python-lzo # lzo needs the C library: brew install lzo
curl -L -o fa.jsonl https://kaikki.org/dictionary/Persian/kaikki.org-dictionary-Persian.jsonl
curl -L -o daneshjoo.mdx \
"https://raw.githubusercontent.com/0xdolan/Daneshjoo/main/Daneshjoo%20Dictionary/Daneshjoo%20Dictionary.mdx"
python3 build_dictionary.py
gzip -9 ganjoor-dictionary.db
```
## Why two sources
Neither alone is enough. Measured against every distinct word in five real poems (Hafez ×2,
Saadi's Golestan, Rumi's Masnavi, a Khayyam rubaʿi — 485 words):
| | coverage |
|---|---|
| Daneshjoo alone | 71% |
| Wiktionary alone + its form index | ~80% |
| **Both, with the lookup chain** | **88%** |
13 of the 55 remaining misses are Arabic lines quoted inside Persian poems, so Persian coverage
is about 91%.
Wiktionary's export is 93 MB, of which the definitions are 0.94 MB — the rest is inflection
tables, etymology templates, IPA and descendants. The build keeps the definitions and the
form→lemma index (149,589 pairs) and drops the rest. That index is what resolves conjugated
verbs: `افتاد → افتادن`, `بگشاید → گشودن`, `دانند → دانستن`.
## Licences
Each row carries its `source`, so attribution stays accurate and either source can be dropped
without rebuilding the other.
- **Wiktionary** — CC BY-SA 3.0. The generated database is therefore also CC BY-SA 3.0.
- **Daneshjoo** — the repository states MIT. Note the underlying lexicon is a published Iranian
dictionary, so that relicensing is worth verifying before relying on it.