Add the Urdu dictionaries, and suggest near words when nothing matches

Two Urdu sources join the Persian ones. Wiktionary's Urdu extract gives Urdu
headwords glossed in English, and Urdu Wiktionary itself gives definitions
written in Urdu — the only source here that does. The latter is thin, about
3,100 usable entries out of 31,000 pages since many are stubs, but for a word
it carries an Urdu reader is better served by it than by a translation into
English: عشق comes back as شدید جذبۂ محبت، گہری چاہت، محبت، پریم، پیار.

Every definition now names the dictionary and its language pair, and lays out
in the direction its own script reads, so an Urdu definition is right-aligned
beside a left-aligned English one.

When nothing matches, the sheet offers near words ranked by how many letters
they share with what was looked up, drawn from an index range scan on the
leading letters rather than a scan of the whole table. خودکامی, which has no
entry, offers خودکامه — the lemma it wants.

Measured honestly: the Urdu sources add little coverage over Persian — nine
words from the English-glossed extract, two from Urdu Wiktionary, against
1,285 from thirteen poems. They are here because an Urdu reader wants Urdu,
not because they widen the net.

The suggestion ranking is verified against the built database rather than only
on device: خودکامی → خودکامه, شیرازی → شیراز, مشکلها → مشکل. On an API 36
emulator all four sources answer عشق with their labels.

Known rough edge: affix stripping across four languages can mislead. ناولها
reaches ناول, the Urdu for "novel", which is not what Hafez meant. The matched
headword is always shown, so it is visible rather than silent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Anas RashidandClaude Opus 5 committed 2026-10-04 17:29:38 +02:00
1 parent 8aad086b08
commit c0cb6a8ab6
11 files changed
+259 -24

No files matched your search

@@ -1,6 +1,7 @@
package com.ganjoor.android
import com.ganjoor.android.data.affixes
import com.ganjoor.android.data.letterOverlap
import com.ganjoor.android.data.normalise
import com.ganjoor.android.data.wordAt
import org.junit.Assert.assertEquals
@@ -106,3 +107,31 @@ class PersianMorphologyTest {
assertTrue(affixes("بها").none { it.length < 2 })
}
}
class LetterOverlapTest {
@Test
fun `an identical word overlaps completely`() {
assertEquals(1f, letterOverlap("عشق", "عشق"), 0.001f)
}
@Test
fun `a suffixed form still scores high against its stem`() {
assertTrue(letterOverlap("مشکل", "مشکلها") > 0.6f)
}
@Test
fun `sharing only a first letter scores low`() {
assertTrue(letterOverlap("عشق", "عبادتگاه") < 0.4f)
}
@Test
fun `letters are counted once each, not by presence alone`() {
// ااا against ا shares one letter, not three
assertEquals(1f / 3f, letterOverlap("ااا", "ا"), 0.001f)
}
@Test
fun `an empty word never matches`() {
assertEquals(0f, letterOverlap("", "عشق"), 0.001f)
}
}