Dictionary coverage, pronunciation, and chapter order #1

Merged
anas merged 5 commits from feature/dictionary-coverage-and-chapter-order into main 2026-10-04 16:12:03 +00:00

5 Commits

Author SHA1 Message Date
Anas Rashid
0d083b6638 Make OLED a switch on the dark themes, and stream recitations
OLED was a sixth theme, which meant choosing it threw away sepia night. It is
a switch now, applied to whichever dark theme is in use: the surfaces move
towards black rather than being replaced by it, so sepia night keeps its warmth
and only the background goes truly black. Text and accents are untouched.
Anyone who had picked the old Black theme comes out as Dark with the switch on.

Recitations: Ganjoor hosts readings of the poems — twenty of Hafez's first
ghazal alone — and the poem screen now plays them. Streamed, never stored,
because a dozen readings of one ghazal would dwarf the poems themselves, which
also makes this the one part of the app that genuinely needs a connection; it
doesn't appear without one. The platform's MediaPlayer rather than ExoPlayer:
one URL, play and pause, nothing to justify a media library.

Pronunciations were laid out as a flowing row, which put Latin IPA and Urdu
spelling in the same right-to-left run — the chips reordered and each label
ended up under someone else's value. They are one per line now, in a single
direction.

Verified on an API 36 emulator: sepia night with the switch on is black with
warm cards; the player shows play on arrival with no network call, and streams
on press.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 18:05:21 +02:00
Anas Rashid
0df736fd89 Show how a word is pronounced, Classical Persian first
Wiktionary carries pronunciation for 45,000 of these headwords and the build
was throwing it away. It is now kept: 101,306 entries of IPA, each tagged with
the variety it belongs to, which matters here because a word in a 14th-century
ghazal was not said the way Tehran says it now — عشق is /ˈʔiʃq/ in Classical
Persian and [ʔeʃɢ̥] in Iran today, and the app leads with the former.

The IPA's dots and stress marks are the syllable breakdown. Urdu Wiktionary
adds its own, in Urdu script: the fully vowelled spelling عِشْق and the split
عِش + قوں, both pulled out of its wikitext.

Costs 7.7 MB of database, taking it to 94 MB.

Verified on an API 36 emulator: tapping عشق shows the Classical Persian IPA,
the Urdu syllable split, the vowelled spelling and two modern variants above
the definitions.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 17:48:12 +02:00
Anas Rashid
b16d020ce0 Add the full Arabic Wiktionary, and answer in Urdu before English
Classical Persian quotes Arabic outright — Hafez opens with a whole hemistich
of it — and none of it resolved before. The full Arabic export brings 36,627
entries and 819,608 new form pairs, and the forms are the point: السّاقی,
الناس, تَلْقَ and تَهْوی are all conjugated or carry the article, so they only
reach a definition through that index. Nine of the ten Arabic words in the
sample now answer where none did.

The cost is real and worth stating: the database goes from 28 MB to 87 MB and
the release APK from 10.6 MB to 29.2 MB, with another 87 MB unpacked on first
run, so about 116 MB installed.

Results are now ordered the way a reader of this app wants them: the Urdu
definition first, because it needs no translating, then Persian, then sources
keyed on another language, with English arriving only through whatever is left.
Arabic sorts last outright — its index is larger than every other source
combined, which makes it the likeliest to match by coincidence.

No Persian-to-Urdu dictionary. Wiktionary's Persian entries carry no
translations at all; the tables live only on English pages, in a 3.3 GB export,
so it would mean pivoting through an English sense. A sample of that file
projects about 6,900 Persian words with any Urdu equivalent, mostly modern
dictionary vocabulary rather than the language of the poems — not worth the
pivot. tools/README.md records why, so the question doesn't get reopened from
scratch.

Verified on an API 36 emulator: عشق returns all five sources with the Urdu
definition at the top.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 17:42:09 +02:00
Anas Rashid
c0cb6a8ab6 Add the Urdu dictionaries, and suggest near words when nothing matches
Two Urdu sources join the Persian ones. Wiktionary's Urdu extract gives Urdu
headwords glossed in English, and Urdu Wiktionary itself gives definitions
written in Urdu — the only source here that does. The latter is thin, about
3,100 usable entries out of 31,000 pages since many are stubs, but for a word
it carries an Urdu reader is better served by it than by a translation into
English: عشق comes back as شدید جذبۂ محبت، گہری چاہت، محبت، پریم، پیار.

Every definition now names the dictionary and its language pair, and lays out
in the direction its own script reads, so an Urdu definition is right-aligned
beside a left-aligned English one.

When nothing matches, the sheet offers near words ranked by how many letters
they share with what was looked up, drawn from an index range scan on the
leading letters rather than a scan of the whole table. خودکامی, which has no
entry, offers خودکامه — the lemma it wants.

Measured honestly: the Urdu sources add little coverage over Persian — nine
words from the English-glossed extract, two from Urdu Wiktionary, against
1,285 from thirteen poems. They are here because an Urdu reader wants Urdu,
not because they widen the net.

The suggestion ranking is verified against the built database rather than only
on device: خودکامی → خودکامه, شیرازی → شیراز, مشکلها → مشکل. On an API 36
emulator all four sources answer عشق with their labels.

Known rough edge: affix stripping across four languages can mislead. ناولها
reaches ناول, the Urdu for "novel", which is not what Hafez meant. The matched
headword is always shown, so it is visible rather than silent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 17:29:38 +02:00
Anas Rashid
8aad086b08 Fix chapter order, and cover more words without more data
Chapter order: Golestan's دیباچه was rendering below the eight باب because the
listing drew every child category and then every poem. ganjoor.net interleaves
them — a book's preface first, then its chapters, then its remaining poems, so
Hafez reads مقدّمه, five collections, مثنوی, ساقی‌نامه. Ganjoor decides this
with each poem's MixedModeOrder (1 above the chapters, 0 below), which the
exported _cat.json doesn't carry and the live API only exposes one poem at a
time, so prefaces are matched by title for now. Adding MixedModeOrder to the
Poems entries in ganjoor-data would make it exact.

Coverage: I went looking for Arabic and Urdu Wiktionary and measured them
instead of assuming. Against 1,285 distinct words from thirteen poems, Urdu
added nine words and Arabic could reach at most five — the uncovered words
were never Arabic, they were Persian morphology the lookup didn't handle:
enclitic pronouns (آیدت, باشدش, تربتش), negation stacked on a prefix
(برنیاید), and compounds. Deepening the affix chain to two passes takes 87% to
91% with no new data at all, so neither dictionary ships.

Each definition now names its language pair rather than just its source, so a
reader can tell what they are looking at.

Dropped the selection-toolbar lookup: Compose 1.10 stopped routing
SelectionContainer through LocalTextToolbar, so a custom toolbar is never
asked to show, and the replacement in foundation's contextmenu package is
internal. Tapping a word already does the lookup; the note in PoemScreen says
when to revisit.

Verified on an API 36 emulator: Golestan lists دیباچه first, and tapping
نافه‌ای resolves through its affix to ناف with both sources labelled.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 17:20:24 +02:00