Commit Graph

7 Commits

Author SHA1 Message Date
Anas Rashid
b16d020ce0 Add the full Arabic Wiktionary, and answer in Urdu before English
Classical Persian quotes Arabic outright — Hafez opens with a whole hemistich
of it — and none of it resolved before. The full Arabic export brings 36,627
entries and 819,608 new form pairs, and the forms are the point: السّاقی,
الناس, تَلْقَ and تَهْوی are all conjugated or carry the article, so they only
reach a definition through that index. Nine of the ten Arabic words in the
sample now answer where none did.

The cost is real and worth stating: the database goes from 28 MB to 87 MB and
the release APK from 10.6 MB to 29.2 MB, with another 87 MB unpacked on first
run, so about 116 MB installed.

Results are now ordered the way a reader of this app wants them: the Urdu
definition first, because it needs no translating, then Persian, then sources
keyed on another language, with English arriving only through whatever is left.
Arabic sorts last outright — its index is larger than every other source
combined, which makes it the likeliest to match by coincidence.

No Persian-to-Urdu dictionary. Wiktionary's Persian entries carry no
translations at all; the tables live only on English pages, in a 3.3 GB export,
so it would mean pivoting through an English sense. A sample of that file
projects about 6,900 Persian words with any Urdu equivalent, mostly modern
dictionary vocabulary rather than the language of the poems — not worth the
pivot. tools/README.md records why, so the question doesn't get reopened from
scratch.

Verified on an API 36 emulator: عشق returns all five sources with the Urdu
definition at the top.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 17:42:09 +02:00
Anas Rashid
c0cb6a8ab6 Add the Urdu dictionaries, and suggest near words when nothing matches
Two Urdu sources join the Persian ones. Wiktionary's Urdu extract gives Urdu
headwords glossed in English, and Urdu Wiktionary itself gives definitions
written in Urdu — the only source here that does. The latter is thin, about
3,100 usable entries out of 31,000 pages since many are stubs, but for a word
it carries an Urdu reader is better served by it than by a translation into
English: عشق comes back as شدید جذبۂ محبت، گہری چاہت، محبت، پریم، پیار.

Every definition now names the dictionary and its language pair, and lays out
in the direction its own script reads, so an Urdu definition is right-aligned
beside a left-aligned English one.

When nothing matches, the sheet offers near words ranked by how many letters
they share with what was looked up, drawn from an index range scan on the
leading letters rather than a scan of the whole table. خودکامی, which has no
entry, offers خودکامه — the lemma it wants.

Measured honestly: the Urdu sources add little coverage over Persian — nine
words from the English-glossed extract, two from Urdu Wiktionary, against
1,285 from thirteen poems. They are here because an Urdu reader wants Urdu,
not because they widen the net.

The suggestion ranking is verified against the built database rather than only
on device: خودکامی → خودکامه, شیرازی → شیراز, مشکلها → مشکل. On an API 36
emulator all four sources answer عشق with their labels.

Known rough edge: affix stripping across four languages can mislead. ناولها
reaches ناول, the Urdu for "novel", which is not what Hafez meant. The matched
headword is always shown, so it is visible rather than silent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 17:29:38 +02:00
Anas Rashid
8aad086b08 Fix chapter order, and cover more words without more data
Chapter order: Golestan's دیباچه was rendering below the eight باب because the
listing drew every child category and then every poem. ganjoor.net interleaves
them — a book's preface first, then its chapters, then its remaining poems, so
Hafez reads مقدّمه, five collections, مثنوی, ساقی‌نامه. Ganjoor decides this
with each poem's MixedModeOrder (1 above the chapters, 0 below), which the
exported _cat.json doesn't carry and the live API only exposes one poem at a
time, so prefaces are matched by title for now. Adding MixedModeOrder to the
Poems entries in ganjoor-data would make it exact.

Coverage: I went looking for Arabic and Urdu Wiktionary and measured them
instead of assuming. Against 1,285 distinct words from thirteen poems, Urdu
added nine words and Arabic could reach at most five — the uncovered words
were never Arabic, they were Persian morphology the lookup didn't handle:
enclitic pronouns (آیدت, باشدش, تربتش), negation stacked on a prefix
(برنیاید), and compounds. Deepening the affix chain to two passes takes 87% to
91% with no new data at all, so neither dictionary ships.

Each definition now names its language pair rather than just its source, so a
reader can tell what they are looking at.

Dropped the selection-toolbar lookup: Compose 1.10 stopped routing
SelectionContainer through LocalTextToolbar, so a custom toolbar is never
asked to show, and the replacement in foundation's contextmenu package is
internal. Tapping a word already does the lookup; the note in PoemScreen says
when to revisit.

Verified on an API 36 emulator: Golestan lists دیباچه first, and tapping
نافه‌ای resolves through its affix to ناف with both sources labelled.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 17:20:24 +02:00
Anas Rashid
c548e6ec07 Add a tap-a-word dictionary, layering Wiktionary and Daneshjoo
Tapping a word in a poem now opens its meaning. Neither source covers enough
alone — Daneshjoo reaches 71% of real poem vocabulary — so both ship, each row
carrying its source so the credit stays attached and either can be dropped
later. Together with the lookup chain that reaches 88%, and 13 of the 55
remaining misses are Arabic lines quoted inside Persian poems.

Lookup widens until something matches: the word as written, the lemma it
inflects from, the word with an affix stripped, then the parts of a ZWNJ
compound. The lemma step is what makes classical verse readable, and it comes
from Wiktionary's 149,589 form->lemma pairs: افتاد is only findable as افتادن.
The Arabic definite article is stripped too, since poems quote Arabic.

Two things the build taught me. Wiktionary's 93 MB export is 0.94 MB of
definitions wrapped in inflection tables, etymology templates and IPA, so
tools/build_dictionary.py keeps the definitions and the form index and drops
the rest. And the packager silently gunzips .gz assets and strips the
extension, which shipped the database under a name the code wasn't opening —
every lookup failed silently until the APK listing gave it away.

Hit testing needed care as well: nastaliq is set with 2.4x leading, so most of
a line box is empty space, and a tap there clamps to the line's first
character — which made every tap return the opening word. Taps are now checked
against the baseline band, with the multipliers left as a knob for swapped
fonts. Definitions are English, so they read left-to-right inside the
otherwise right-to-left sheet.

Verified on the release build, where R8 could have broken the SQLite path:
tapping عشق returns both sources, and دانند resolves to دانستن.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 02:02:00 +02:00
Anas Rashid
a5fc608bca Add poem search, an OLED theme, and credit every licence
Search: full-text across the poems, scoped to all poets, a few, or one. The
static data set has no index, so this is ganjoor.net's own search endpoint,
which scopes to one poet at a time — several picked poets run as several
queries and merge. Entry is the poets filter box, which now sits in the bottom
bar within thumb reach and offers its own words to the poem search one tap
further. Snippets show the line the term is actually on.

Themes: a sixth, true black, for OLED panels where an unlit pixel costs no
power. Surfaces step up in near-blacks so cards stay distinguishable.

Licences: every component is now credited in-app under Reading settings >
About & licences, with the full text bundled in assets/licenses — the OFL
requires the licence to travel with the software, so linking to it wasn't
enough. licenses/README.md indexes the same thing for the repo. The credits
list and the licence texts lay out left-to-right in English, since the rest of
the app is right-to-left for the poetry's sake and English prose is unreadable
right-aligned.

Fixes a crash: the settings sheet was composed outside the provider supplying
LocalOpenAbout, so opening it threw.

Search decoding is covered by a unit test against a real API response; the
rest verified on an API 36 emulator.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 01:05:52 +02:00
Anas Rashid
1accb5ee78 Navigate the poem tree instead of the visit history
Breadcrumbs: a poem's path (poet » book » section) is now built from its own
fullTitle and fullUrl and every ancestor is tappable, so the trail is the same
whether the poem was opened by browsing, from the saved list, from a
breadcrumb, or by reading on from the previous poem.

Back: browsing is a tree, so Back climbs it — poem to section to book to poet
to home — rather than retracing however you arrived. Each screen derives its
own parent from its URL and navigation keeps the stack flat, which is what
stops Back from walking forward again into the page you just left. System Back
and the top-bar arrow do the same thing.

The exception is a poem opened from the saved list: that was reached from a
list, not a shelf, so Back returns to the list, and reading on through the
divan keeps that origin.

Plus a Home button on every screen below the poet list.

Verified on an API 36 emulator: drilled Hafez > ghazals > ghazal 3 and walked
back up to the poet list, and opened a saved Khayyam rubai and confirmed Back
returns to the saved list rather than the poet's section.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-04 00:28:52 +02:00
Anas Rashid
abcffbc43b Add offline reading, bookmarks, three UI languages and downloads
Reading
- Naskh is now the default face, and both Arabic-script fonts are registered
  at four weights so the text can be thickened for a lit screen.
- Category listings show each poem's opening line, fetched best-effort from
  api.ganjoor.net since the data set doesn't carry excerpts.
- Tap a couplet to save that passage or copy it; a saved passage keeps a
  tappable link back to its poem. Poem text is selectable for plain copying.

Languages
- Persian, Urdu and English, Persian by default whatever the phone's locale,
  applied in attachBaseContext and switchable from the top bar or the sheet.
- English chrome uses Libron (OFL); poems stay naskh or nastaliq throughout.
- Each language gets its own values-* folder: with Farsi only in values/,
  Android was resolving it to the Urdu strings, since a same-script locale
  outranks the default.

Offline
- Downloaded poets mirror the data set's layout under filesDir, so offline
  mode is one lookup rather than a parallel path. Downloads run one poet at
  a time, skip what's on disk, and resume by restarting.
- Downloads screen lists every poet with their portrait and tick boxes for
  picking several, plus storage used, delete, and download-everything behind
  a size warning.
- Offline mode refuses the network and says what's missing rather than
  blaming the connection.

Also: poet sort (Ganjoor's order or alphabetical, via a Persian collator),
RTL nav transitions, an original shamsa launcher icon with a monochrome layer
for themed icons, and preferences written with commit so they survive the
process being killed.

F-Droid: dependenciesInfo off, optional signing so the build works with no
keystore, fastlane metadata in three languages, cleartext traffic disabled.

Verified on an API 36 emulator: downloaded a poet, pulled the emulator's
network, read them offline, and confirmed an undownloaded poet reports being
missing rather than erroring.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-03 23:20:42 +02:00