Add the Urdu dictionaries, and suggest near words when nothing matches
Two Urdu sources join the Persian ones. Wiktionary's Urdu extract gives Urdu headwords glossed in English, and Urdu Wiktionary itself gives definitions written in Urdu — the only source here that does. The latter is thin, about 3,100 usable entries out of 31,000 pages since many are stubs, but for a word it carries an Urdu reader is better served by it than by a translation into English: عشق comes back as شدید جذبۂ محبت، گہری چاہت، محبت، پریم، پیار. Every definition now names the dictionary and its language pair, and lays out in the direction its own script reads, so an Urdu definition is right-aligned beside a left-aligned English one. When nothing matches, the sheet offers near words ranked by how many letters they share with what was looked up, drawn from an index range scan on the leading letters rather than a scan of the whole table. خودکامی, which has no entry, offers خودکامه — the lemma it wants. Measured honestly: the Urdu sources add little coverage over Persian — nine words from the English-glossed extract, two from Urdu Wiktionary, against 1,285 from thirteen poems. They are here because an Urdu reader wants Urdu, not because they widen the net. The suggestion ranking is verified against the built database rather than only on device: خودکامی → خودکامه, شیرازی → شیراز, مشکلها → مشکل. On an API 36 emulator all four sources answer عشق with their labels. Known rough edge: affix stripping across four languages can mislead. ناولها reaches ناول, the Urdu for "novel", which is not what Hafez meant. The matched headword is always shown, so it is visible rather than silent. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
1 parent
8aad086b08
commit
c0cb6a8ab6
11 files changed
+259
-24
No files matched your search
@@ -1,6 +1,7 @@
|
||||
package com.ganjoor.android
|
||||
|
||||
import com.ganjoor.android.data.affixes
|
||||
import com.ganjoor.android.data.letterOverlap
|
||||
import com.ganjoor.android.data.normalise
|
||||
import com.ganjoor.android.data.wordAt
|
||||
import org.junit.Assert.assertEquals
|
||||
@@ -106,3 +107,31 @@ class PersianMorphologyTest {
|
||||
assertTrue(affixes("بها").none { it.length < 2 })
|
||||
}
|
||||
}
|
||||
|
||||
class LetterOverlapTest {
|
||||
@Test
|
||||
fun `an identical word overlaps completely`() {
|
||||
assertEquals(1f, letterOverlap("عشق", "عشق"), 0.001f)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun `a suffixed form still scores high against its stem`() {
|
||||
assertTrue(letterOverlap("مشکل", "مشکلها") > 0.6f)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun `sharing only a first letter scores low`() {
|
||||
assertTrue(letterOverlap("عشق", "عبادتگاه") < 0.4f)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun `letters are counted once each, not by presence alone`() {
|
||||
// ااا against ا shares one letter, not three
|
||||
assertEquals(1f / 3f, letterOverlap("ااا", "ا"), 0.001f)
|
||||
}
|
||||
|
||||
@Test
|
||||
fun `an empty word never matches`() {
|
||||
assertEquals(0f, letterOverlap("", "عشق"), 0.001f)
|
||||
}
|
||||
}
|
||||
Reference in new issue
Block a user