Add a tap-a-word dictionary, layering Wiktionary and Daneshjoo
Tapping a word in a poem now opens its meaning. Neither source covers enough alone — Daneshjoo reaches 71% of real poem vocabulary — so both ship, each row carrying its source so the credit stays attached and either can be dropped later. Together with the lookup chain that reaches 88%, and 13 of the 55 remaining misses are Arabic lines quoted inside Persian poems. Lookup widens until something matches: the word as written, the lemma it inflects from, the word with an affix stripped, then the parts of a ZWNJ compound. The lemma step is what makes classical verse readable, and it comes from Wiktionary's 149,589 form->lemma pairs: افتاد is only findable as افتادن. The Arabic definite article is stripped too, since poems quote Arabic. Two things the build taught me. Wiktionary's 93 MB export is 0.94 MB of definitions wrapped in inflection tables, etymology templates and IPA, so tools/build_dictionary.py keeps the definitions and the form index and drops the rest. And the packager silently gunzips .gz assets and strips the extension, which shipped the database under a name the code wasn't opening — every lookup failed silently until the APK listing gave it away. Hit testing needed care as well: nastaliq is set with 2.4x leading, so most of a line box is empty space, and a tap there clamps to the line's first character — which made every tap return the opening word. Taps are now checked against the baseline band, with the multipliers left as a knob for swapped fonts. Definitions are English, so they read left-to-right inside the otherwise right-to-left sheet. Verified on the release build, where R8 could have broken the SQLite path: tapping عشق returns both sources, and دانند resolves to دانستن. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
b545220bbc
commit
c548e6ec07
21
README.md
21
README.md
@ -58,6 +58,27 @@ downloaded, rather than blaming your connection.
|
|||||||
Downloads run one poet at a time and are resumable: anything already on disk is skipped, so
|
Downloads run one poet at a time and are resumable: anything already on disk is skipped, so
|
||||||
restarting an interrupted download picks up where it left off.
|
restarting an interrupted download picks up where it left off.
|
||||||
|
|
||||||
|
## Dictionary
|
||||||
|
|
||||||
|
Tap any word in a poem for its meaning. The bundled database layers two sources — neither is
|
||||||
|
enough alone:
|
||||||
|
|
||||||
|
| | coverage of real poem vocabulary |
|
||||||
|
|---|---|
|
||||||
|
| Daneshjoo alone | 71% |
|
||||||
|
| **Both, with the lookup chain** | **88%** |
|
||||||
|
|
||||||
|
Measured over every distinct word in five poems (Hafez ×2, Golestan, Masnavi, a Khayyam rubaʿi).
|
||||||
|
13 of the 55 remaining misses are Arabic lines quoted inside Persian poems, so Persian coverage
|
||||||
|
is about 91%.
|
||||||
|
|
||||||
|
Lookup widens until something matches: the word as written, then the lemma it inflects from,
|
||||||
|
then with an affix stripped, then the parts of a ZWNJ compound. The lemma step is what makes
|
||||||
|
classical verse readable — `افتاد` is only findable as `افتادن`, and Wiktionary ships 149,589
|
||||||
|
form→lemma pairs that make that possible.
|
||||||
|
|
||||||
|
See [`tools/README.md`](tools/README.md) to rebuild it, and for the licensing of each source.
|
||||||
|
|
||||||
## Where the poems come from
|
## Where the poems come from
|
||||||
|
|
||||||
There is no backend. Every "endpoint" is a JSON file in
|
There is no backend. Every "endpoint" is a JSON file in
|
||||||
|
|||||||
BIN
app/src/main/assets/dictionary.db
Normal file
BIN
app/src/main/assets/dictionary.db
Normal file
Binary file not shown.
@ -12,6 +12,7 @@ import androidx.compose.ui.platform.LocalLayoutDirection
|
|||||||
import androidx.compose.ui.unit.LayoutDirection
|
import androidx.compose.ui.unit.LayoutDirection
|
||||||
import androidx.core.graphics.drawable.toDrawable
|
import androidx.core.graphics.drawable.toDrawable
|
||||||
import com.ganjoor.android.data.Bookmarks
|
import com.ganjoor.android.data.Bookmarks
|
||||||
|
import com.ganjoor.android.data.Dictionary
|
||||||
import com.ganjoor.android.data.Ganjoor
|
import com.ganjoor.android.data.Ganjoor
|
||||||
import com.ganjoor.android.data.LocalBookmarks
|
import com.ganjoor.android.data.LocalBookmarks
|
||||||
import com.ganjoor.android.ui.GanjoorApp
|
import com.ganjoor.android.ui.GanjoorApp
|
||||||
@ -37,6 +38,7 @@ class MainActivity : ComponentActivity() {
|
|||||||
override fun onCreate(savedInstanceState: Bundle?) {
|
override fun onCreate(savedInstanceState: Bundle?) {
|
||||||
super.onCreate(savedInstanceState)
|
super.onCreate(savedInstanceState)
|
||||||
Ganjoor.init(applicationContext)
|
Ganjoor.init(applicationContext)
|
||||||
|
Dictionary.init(applicationContext)
|
||||||
|
|
||||||
// Before the first frame: otherwise the window keeps the platform's white through
|
// Before the first frame: otherwise the window keeps the platform's white through
|
||||||
// startup and every screen transition, whatever theme is chosen.
|
// startup and every screen transition, whatever theme is chosen.
|
||||||
|
|||||||
140
app/src/main/java/com/ganjoor/android/data/Dictionary.kt
Normal file
140
app/src/main/java/com/ganjoor/android/data/Dictionary.kt
Normal file
@ -0,0 +1,140 @@
|
|||||||
|
package com.ganjoor.android.data
|
||||||
|
|
||||||
|
import android.content.Context
|
||||||
|
import android.database.sqlite.SQLiteDatabase
|
||||||
|
import kotlinx.coroutines.Dispatchers
|
||||||
|
import kotlinx.coroutines.withContext
|
||||||
|
import java.io.File
|
||||||
|
import java.text.Normalizer
|
||||||
|
|
||||||
|
|
||||||
|
/** One definition, and where it came from, so the credit stays attached to the text. */
|
||||||
|
data class Definition(val word: String, val gloss: String, val source: String)
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Word lookup over a bundled SQLite built from Wiktionary (CC BY-SA 3.0) and Daneshjoo.
|
||||||
|
* See `tools/build_dictionary.py`; the asset ships gzipped and is unpacked once on first use.
|
||||||
|
*/
|
||||||
|
object Dictionary {
|
||||||
|
/** Guards against a copy interrupted half-way leaving an unopenable file behind. */
|
||||||
|
private const val ASSET_BYTES = 22_298_624L
|
||||||
|
|
||||||
|
// Not a .gz: the build packager silently gunzips those and drops the extension, which left
|
||||||
|
// the asset under a different name than the code was opening.
|
||||||
|
private const val ASSET = "dictionary.db"
|
||||||
|
private lateinit var appContext: Context
|
||||||
|
|
||||||
|
@Volatile
|
||||||
|
private var db: SQLiteDatabase? = null
|
||||||
|
|
||||||
|
fun init(context: Context) {
|
||||||
|
appContext = context.applicationContext
|
||||||
|
}
|
||||||
|
|
||||||
|
private fun open(): SQLiteDatabase? {
|
||||||
|
db?.let { return it }
|
||||||
|
return synchronized(this) {
|
||||||
|
db ?: runCatching {
|
||||||
|
val file = File(appContext.filesDir, "dictionary.db")
|
||||||
|
if (file.length() != ASSET_BYTES) {
|
||||||
|
appContext.assets.open(ASSET).use { input ->
|
||||||
|
file.outputStream().use { input.copyTo(it) }
|
||||||
|
}
|
||||||
|
}
|
||||||
|
SQLiteDatabase.openDatabase(file.path, null, SQLiteDatabase.OPEN_READONLY)
|
||||||
|
}.getOrNull()?.also { db = it }
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Looks a word up, widening the search until something matches:
|
||||||
|
*
|
||||||
|
* 1. the word as written;
|
||||||
|
* 2. the lemma it inflects from — Persian verbs conjugate heavily, and `افتاد` is only
|
||||||
|
* findable as `افتادن`;
|
||||||
|
* 3. the word with a common prefix or suffix removed;
|
||||||
|
* 4. the parts of a ZWNJ compound, so `بیروزی` finds `روزی`.
|
||||||
|
*/
|
||||||
|
suspend fun lookup(raw: String): List<Definition> = withContext(Dispatchers.IO) {
|
||||||
|
val database = open() ?: return@withContext emptyList()
|
||||||
|
val word = normalise(raw)
|
||||||
|
if (word.isEmpty()) return@withContext emptyList()
|
||||||
|
|
||||||
|
direct(database, word)
|
||||||
|
.ifEmpty { lemmas(database, word).flatMap { direct(database, it) } }
|
||||||
|
.ifEmpty { affixes(word).firstNotNullOfOrNull { direct(database, it).ifEmpty { null } }.orEmpty() }
|
||||||
|
.ifEmpty {
|
||||||
|
normalise(raw, keepZwnj = true).split(ZWNJ)
|
||||||
|
.filter { it.length > 1 }
|
||||||
|
.firstNotNullOfOrNull { direct(database, normalise(it)).ifEmpty { null } }
|
||||||
|
.orEmpty()
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private fun direct(database: SQLiteDatabase, word: String): List<Definition> =
|
||||||
|
database.rawQuery(
|
||||||
|
"SELECT display, gloss, source FROM entry WHERE word = ? LIMIT 12",
|
||||||
|
arrayOf(word),
|
||||||
|
).use { cursor ->
|
||||||
|
buildList {
|
||||||
|
while (cursor.moveToNext()) {
|
||||||
|
add(Definition(cursor.getString(0), cursor.getString(1), cursor.getString(2)))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private fun lemmas(database: SQLiteDatabase, form: String): List<String> =
|
||||||
|
database.rawQuery(
|
||||||
|
"SELECT lemma FROM form WHERE form = ? LIMIT 6",
|
||||||
|
arrayOf(form),
|
||||||
|
).use { cursor ->
|
||||||
|
buildList { while (cursor.moveToNext()) add(cursor.getString(0)) }
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private const val ZWNJ = ''
|
||||||
|
|
||||||
|
private val SUFFIXES = listOf("ها", "اش", "ش", "م", "ت", "را", "ی", "ان")
|
||||||
|
// "ال" is the Arabic definite article: poems quote Arabic, so السّاقی has to reach ساقی.
|
||||||
|
private val PREFIXES = listOf("ال", "می", "بر", "ب")
|
||||||
|
|
||||||
|
/** Candidate stems after stripping one common affix. Order matters: longest affix first. */
|
||||||
|
internal fun affixes(word: String): List<String> = buildList {
|
||||||
|
SUFFIXES.forEach { if (word.endsWith(it) && word.length > it.length + 1) add(word.dropLast(it.length)) }
|
||||||
|
PREFIXES.forEach { if (word.startsWith(it) && word.length > it.length + 1) add(word.drop(it.length)) }
|
||||||
|
}
|
||||||
|
|
||||||
|
private val HARAKAT = (0x064B..0x0652) + listOf(0x0670, 0x0640) + (0x0610..0x0615)
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Folds a word to the form the dictionary is keyed by: diacritics dropped, and the Arabic
|
||||||
|
* letters Persian writes differently (ي ك ة) folded to their Persian shapes, because poems and
|
||||||
|
* dictionary headwords disagree about them constantly.
|
||||||
|
*/
|
||||||
|
internal fun normalise(text: String, keepZwnj: Boolean = false): String {
|
||||||
|
val stripped = Normalizer.normalize(text, Normalizer.Form.NFD)
|
||||||
|
.filter { it.code !in HARAKAT }
|
||||||
|
val folded = Normalizer.normalize(stripped, Normalizer.Form.NFC).map { char ->
|
||||||
|
when (char) {
|
||||||
|
'ي', 'ى' -> 'ی'
|
||||||
|
'ك' -> 'ک'
|
||||||
|
'ة' -> 'ه'
|
||||||
|
else -> char
|
||||||
|
}
|
||||||
|
}.joinToString("")
|
||||||
|
return (if (keepZwnj) folded else folded.replace(ZWNJ.toString(), "")).trim()
|
||||||
|
}
|
||||||
|
|
||||||
|
/** The whole word surrounding [index], for turning a tap into something to look up. */
|
||||||
|
internal fun wordAt(text: String, index: Int): String? {
|
||||||
|
if (text.isEmpty()) return null
|
||||||
|
val at = index.coerceIn(0, text.length - 1)
|
||||||
|
if (!text[at].isWordChar()) return null
|
||||||
|
var start = at
|
||||||
|
while (start > 0 && text[start - 1].isWordChar()) start--
|
||||||
|
var end = at
|
||||||
|
while (end < text.length - 1 && text[end + 1].isWordChar()) end++
|
||||||
|
return text.substring(start, end + 1).trim(ZWNJ).takeIf { it.length > 1 }
|
||||||
|
}
|
||||||
|
|
||||||
|
private fun Char.isWordChar() = this in ''..'ۿ' || this == ZWNJ
|
||||||
@ -62,6 +62,19 @@ private val CREDITS = listOf(
|
|||||||
"No licence stated by the publisher",
|
"No licence stated by the publisher",
|
||||||
url = "https://github.com/anas-rashid/ganjoor-data",
|
url = "https://github.com/anas-rashid/ganjoor-data",
|
||||||
),
|
),
|
||||||
|
Credit(
|
||||||
|
"Wiktionary",
|
||||||
|
"Wiktionary contributors — the word definitions, and the inflected-form index that " +
|
||||||
|
"finds a conjugated verb's dictionary entry",
|
||||||
|
"CC BY-SA 3.0. The bundled dictionary is therefore also CC BY-SA 3.0.",
|
||||||
|
url = "https://en.wiktionary.org",
|
||||||
|
),
|
||||||
|
Credit(
|
||||||
|
"Daneshjoo Dictionary",
|
||||||
|
"Layered under Wiktionary for the words it doesn't carry",
|
||||||
|
"The source repository states MIT",
|
||||||
|
url = "https://github.com/0xdolan/Daneshjoo",
|
||||||
|
),
|
||||||
Credit(
|
Credit(
|
||||||
"Noto Naskh Arabic",
|
"Noto Naskh Arabic",
|
||||||
"The Noto Project Authors",
|
"The Noto Project Authors",
|
||||||
|
|||||||
@ -2,6 +2,7 @@ package com.ganjoor.android.ui
|
|||||||
|
|
||||||
import androidx.activity.compose.BackHandler
|
import androidx.activity.compose.BackHandler
|
||||||
import androidx.compose.foundation.clickable
|
import androidx.compose.foundation.clickable
|
||||||
|
import androidx.compose.foundation.gestures.detectTapGestures
|
||||||
import androidx.compose.foundation.layout.Arrangement
|
import androidx.compose.foundation.layout.Arrangement
|
||||||
import androidx.compose.foundation.layout.FlowRow
|
import androidx.compose.foundation.layout.FlowRow
|
||||||
import androidx.compose.foundation.layout.Column
|
import androidx.compose.foundation.layout.Column
|
||||||
@ -34,7 +35,11 @@ import androidx.compose.runtime.mutableStateOf
|
|||||||
import androidx.compose.runtime.remember
|
import androidx.compose.runtime.remember
|
||||||
import androidx.compose.runtime.setValue
|
import androidx.compose.runtime.setValue
|
||||||
import androidx.compose.ui.Modifier
|
import androidx.compose.ui.Modifier
|
||||||
|
import androidx.compose.ui.geometry.Offset
|
||||||
|
import androidx.compose.ui.input.pointer.pointerInput
|
||||||
|
import androidx.compose.ui.text.TextLayoutResult
|
||||||
import androidx.compose.ui.platform.LocalClipboardManager
|
import androidx.compose.ui.platform.LocalClipboardManager
|
||||||
|
import androidx.compose.ui.platform.LocalDensity
|
||||||
import androidx.compose.ui.res.stringResource
|
import androidx.compose.ui.res.stringResource
|
||||||
import androidx.compose.ui.text.AnnotatedString
|
import androidx.compose.ui.text.AnnotatedString
|
||||||
import androidx.compose.ui.text.style.TextAlign
|
import androidx.compose.ui.text.style.TextAlign
|
||||||
@ -43,6 +48,7 @@ import androidx.compose.ui.unit.dp
|
|||||||
import com.ganjoor.android.R
|
import com.ganjoor.android.R
|
||||||
import com.ganjoor.android.data.Bookmark
|
import com.ganjoor.android.data.Bookmark
|
||||||
import com.ganjoor.android.data.breadcrumbs
|
import com.ganjoor.android.data.breadcrumbs
|
||||||
|
import com.ganjoor.android.data.wordAt
|
||||||
import com.ganjoor.android.data.Ganjoor
|
import com.ganjoor.android.data.Ganjoor
|
||||||
import com.ganjoor.android.data.LocalBookmarks
|
import com.ganjoor.android.data.LocalBookmarks
|
||||||
import com.ganjoor.android.data.PoemRef
|
import com.ganjoor.android.data.PoemRef
|
||||||
@ -61,6 +67,8 @@ fun PoemScreen(
|
|||||||
) {
|
) {
|
||||||
BackHandler(onBack = onUp)
|
BackHandler(onBack = onUp)
|
||||||
|
|
||||||
|
var tappedWord by remember { mutableStateOf<String?>(null) }
|
||||||
|
|
||||||
Load(key = fullUrl, block = { Ganjoor.poem(fullUrl) }) { poem ->
|
Load(key = fullUrl, block = { Ganjoor.poem(fullUrl) }) { poem ->
|
||||||
val prefs = LocalSettings.current.value
|
val prefs = LocalSettings.current.value
|
||||||
val style = readingStyle(prefs.font, prefs.fontSize, prefs.fontWeight.weight)
|
val style = readingStyle(prefs.font, prefs.fontSize, prefs.fontWeight.weight)
|
||||||
@ -128,6 +136,7 @@ fun PoemScreen(
|
|||||||
style = style,
|
style = style,
|
||||||
showSummaries = prefs.showSummaries,
|
showSummaries = prefs.showSummaries,
|
||||||
source = Bookmark(fullUrl, poem.title, poem.fullTitle),
|
source = Bookmark(fullUrl, poem.title, poem.fullTitle),
|
||||||
|
onWord = { tappedWord = it },
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
|
|
||||||
@ -159,6 +168,10 @@ fun PoemScreen(
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
tappedWord?.let { word ->
|
||||||
|
WordSheet(word = word, onDismiss = { tappedWord = null })
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/** The poem's path, with every ancestor tappable: poet » book » section » this poem. */
|
/** The poem's path, with every ancestor tappable: poet » book » section » this poem. */
|
||||||
@ -216,28 +229,19 @@ private fun Couplet(
|
|||||||
style: androidx.compose.ui.text.TextStyle,
|
style: androidx.compose.ui.text.TextStyle,
|
||||||
showSummaries: Boolean,
|
showSummaries: Boolean,
|
||||||
source: Bookmark,
|
source: Bookmark,
|
||||||
|
onWord: (String) -> Unit,
|
||||||
) {
|
) {
|
||||||
// Tap, not long-press: long-press belongs to the text selection this sits inside.
|
// Tap, not long-press: long-press belongs to the text selection this sits inside.
|
||||||
var actionsOpen by remember(couplet) { mutableStateOf(false) }
|
var actionsOpen by remember(couplet) { mutableStateOf(false) }
|
||||||
|
|
||||||
Column(
|
Column(modifier = Modifier.fillMaxWidth().padding(vertical = 6.dp)) {
|
||||||
modifier = Modifier
|
|
||||||
.fillMaxWidth()
|
|
||||||
.clickable { actionsOpen = !actionsOpen }
|
|
||||||
.padding(vertical = 6.dp)
|
|
||||||
) {
|
|
||||||
couplet.forEach { verse ->
|
couplet.forEach { verse ->
|
||||||
Text(
|
VerseText(
|
||||||
text = verse.text,
|
verse = verse,
|
||||||
style = style,
|
style = style,
|
||||||
textAlign = when (verse.position) {
|
onWord = onWord,
|
||||||
Verse.RIGHT -> TextAlign.Start
|
// A tap that lands between words still opens the couplet's own actions.
|
||||||
Verse.LEFT -> TextAlign.End
|
onElsewhere = { actionsOpen = !actionsOpen },
|
||||||
Verse.CENTERED_1, Verse.CENTERED_2 -> TextAlign.Center
|
|
||||||
// Single / Paragraph / Comment: prose, so let it fill the column.
|
|
||||||
else -> TextAlign.Justify
|
|
||||||
},
|
|
||||||
modifier = Modifier.fillMaxWidth(),
|
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
if (showSummaries) {
|
if (showSummaries) {
|
||||||
@ -258,6 +262,67 @@ private fun Couplet(
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* One hemistich. Tapping a word looks it up; tapping between words falls through to the
|
||||||
|
* couplet's save and copy actions, so both live on the same gesture without fighting.
|
||||||
|
*/
|
||||||
|
@Composable
|
||||||
|
private fun VerseText(
|
||||||
|
verse: Verse,
|
||||||
|
style: androidx.compose.ui.text.TextStyle,
|
||||||
|
onWord: (String) -> Unit,
|
||||||
|
onElsewhere: () -> Unit,
|
||||||
|
) {
|
||||||
|
var layout by remember(verse.text) { mutableStateOf<TextLayoutResult?>(null) }
|
||||||
|
val fontSizePx = with(LocalDensity.current) { style.fontSize.toPx() }
|
||||||
|
|
||||||
|
Text(
|
||||||
|
text = verse.text,
|
||||||
|
style = style,
|
||||||
|
textAlign = when (verse.position) {
|
||||||
|
Verse.RIGHT -> TextAlign.Start
|
||||||
|
Verse.LEFT -> TextAlign.End
|
||||||
|
Verse.CENTERED_1, Verse.CENTERED_2 -> TextAlign.Center
|
||||||
|
// Single / Paragraph / Comment: prose, so let it fill the column.
|
||||||
|
else -> TextAlign.Justify
|
||||||
|
},
|
||||||
|
onTextLayout = { layout = it },
|
||||||
|
modifier = Modifier
|
||||||
|
.fillMaxWidth()
|
||||||
|
.pointerInput(verse.text) {
|
||||||
|
detectTapGestures { position ->
|
||||||
|
val word = layout?.let { wordTappedAt(it, verse.text, position, fontSizePx) }
|
||||||
|
if (word != null) onWord(word) else onElsewhere()
|
||||||
|
}
|
||||||
|
},
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* The word actually under [position], or null if the tap missed the glyphs.
|
||||||
|
*
|
||||||
|
* getOffsetForPosition alone isn't enough: nastaliq is set with 2.4x leading, so most of a line
|
||||||
|
* box is empty space above the glyphs, and a tap there clamps to the line's first character —
|
||||||
|
* which made every tap return the opening word. Checking the character's own bounding box is
|
||||||
|
* what distinguishes "on a word" from "in the gap between lines".
|
||||||
|
*/
|
||||||
|
internal fun wordTappedAt(
|
||||||
|
layout: TextLayoutResult,
|
||||||
|
text: String,
|
||||||
|
position: Offset,
|
||||||
|
fontSizePx: Float,
|
||||||
|
): String? {
|
||||||
|
if (text.isEmpty() || fontSizePx <= 0f) return null
|
||||||
|
val offset = layout.getOffsetForPosition(position).coerceIn(0, text.length - 1)
|
||||||
|
val baseline = layout.getLineBaseline(layout.getLineForOffset(offset))
|
||||||
|
// The band the ink actually occupies, measured from the baseline. Nastaliq hangs far above
|
||||||
|
// it and dips a little below; these two multipliers are the knob to turn if a font is
|
||||||
|
// swapped and taps start feeling off.
|
||||||
|
if (position.y < baseline - fontSizePx * 1.4f) return null
|
||||||
|
if (position.y > baseline + fontSizePx * 0.6f) return null
|
||||||
|
return wordAt(text, offset)
|
||||||
|
}
|
||||||
|
|
||||||
/** Save this passage, or copy it. Saving keeps the link back to the poem; copying doesn't. */
|
/** Save this passage, or copy it. Saving keeps the link back to the poem; copying doesn't. */
|
||||||
@Composable
|
@Composable
|
||||||
private fun PassageActions(passage: Bookmark) {
|
private fun PassageActions(passage: Bookmark) {
|
||||||
|
|||||||
106
app/src/main/java/com/ganjoor/android/ui/WordSheet.kt
Normal file
106
app/src/main/java/com/ganjoor/android/ui/WordSheet.kt
Normal file
@ -0,0 +1,106 @@
|
|||||||
|
package com.ganjoor.android.ui
|
||||||
|
|
||||||
|
import androidx.compose.foundation.layout.Arrangement
|
||||||
|
import androidx.compose.foundation.layout.Column
|
||||||
|
import androidx.compose.foundation.layout.fillMaxWidth
|
||||||
|
import androidx.compose.foundation.layout.navigationBarsPadding
|
||||||
|
import androidx.compose.foundation.layout.padding
|
||||||
|
import androidx.compose.foundation.rememberScrollState
|
||||||
|
import androidx.compose.foundation.verticalScroll
|
||||||
|
import androidx.compose.material3.CircularProgressIndicator
|
||||||
|
import androidx.compose.material3.ExperimentalMaterial3Api
|
||||||
|
import androidx.compose.material3.HorizontalDivider
|
||||||
|
import androidx.compose.material3.MaterialTheme
|
||||||
|
import androidx.compose.material3.ModalBottomSheet
|
||||||
|
import androidx.compose.material3.Text
|
||||||
|
import androidx.compose.runtime.Composable
|
||||||
|
import androidx.compose.runtime.CompositionLocalProvider
|
||||||
|
import androidx.compose.runtime.getValue
|
||||||
|
import androidx.compose.runtime.mutableStateOf
|
||||||
|
import androidx.compose.runtime.produceState
|
||||||
|
import androidx.compose.ui.Modifier
|
||||||
|
import androidx.compose.ui.platform.LocalLayoutDirection
|
||||||
|
import androidx.compose.ui.res.stringResource
|
||||||
|
import androidx.compose.ui.unit.LayoutDirection
|
||||||
|
import androidx.compose.ui.unit.dp
|
||||||
|
import com.ganjoor.android.R
|
||||||
|
import com.ganjoor.android.data.Definition
|
||||||
|
import com.ganjoor.android.data.Dictionary
|
||||||
|
import com.ganjoor.android.ui.theme.readingStyle
|
||||||
|
|
||||||
|
/** English prose inside an otherwise right-to-left sheet. */
|
||||||
|
@Composable
|
||||||
|
private fun LeftToRight(content: @Composable () -> Unit) {
|
||||||
|
CompositionLocalProvider(LocalLayoutDirection provides LayoutDirection.Ltr, content = content)
|
||||||
|
}
|
||||||
|
|
||||||
|
/** What the dictionary knows about a tapped word. */
|
||||||
|
@OptIn(ExperimentalMaterial3Api::class)
|
||||||
|
@Composable
|
||||||
|
fun WordSheet(word: String, onDismiss: () -> Unit) {
|
||||||
|
val prefs = LocalSettings.current.value
|
||||||
|
val definitions by produceState<List<Definition>?>(null, word) { value = Dictionary.lookup(word) }
|
||||||
|
|
||||||
|
ModalBottomSheet(onDismissRequest = onDismiss) {
|
||||||
|
Column(
|
||||||
|
modifier = Modifier
|
||||||
|
.navigationBarsPadding()
|
||||||
|
.verticalScroll(rememberScrollState())
|
||||||
|
.padding(horizontal = 20.dp)
|
||||||
|
.padding(bottom = 32.dp),
|
||||||
|
verticalArrangement = Arrangement.spacedBy(12.dp),
|
||||||
|
) {
|
||||||
|
// The headword in the reading font, at reading size: it is a line of poetry, after all.
|
||||||
|
Text(
|
||||||
|
text = word,
|
||||||
|
style = readingStyle(prefs.font, prefs.fontSize, prefs.fontWeight.weight),
|
||||||
|
modifier = Modifier.fillMaxWidth(),
|
||||||
|
)
|
||||||
|
HorizontalDivider()
|
||||||
|
|
||||||
|
when {
|
||||||
|
definitions == null -> CircularProgressIndicator(Modifier.padding(vertical = 16.dp))
|
||||||
|
|
||||||
|
definitions!!.isEmpty() -> LeftToRight {
|
||||||
|
Text(
|
||||||
|
text = stringResource(R.string.no_definition),
|
||||||
|
style = MaterialTheme.typography.bodyMedium,
|
||||||
|
color = MaterialTheme.colorScheme.onSurfaceVariant,
|
||||||
|
modifier = Modifier.fillMaxWidth(),
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
else -> definitions!!.forEach { definition ->
|
||||||
|
Column(Modifier.fillMaxWidth()) {
|
||||||
|
// The headword actually matched, which may be the lemma rather than the
|
||||||
|
// word as it appears in the line. Persian, so it stays right-to-left.
|
||||||
|
if (definition.word != word) {
|
||||||
|
Text(
|
||||||
|
text = definition.word,
|
||||||
|
style = MaterialTheme.typography.titleSmall,
|
||||||
|
color = MaterialTheme.colorScheme.primary,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
// The definitions are English; right-aligning them reads badly.
|
||||||
|
LeftToRight {
|
||||||
|
Column(Modifier.fillMaxWidth()) {
|
||||||
|
Text(definition.gloss, style = MaterialTheme.typography.bodyMedium)
|
||||||
|
Text(
|
||||||
|
text = stringResource(
|
||||||
|
if (definition.source == "wiktionary") {
|
||||||
|
R.string.source_wiktionary
|
||||||
|
} else {
|
||||||
|
R.string.source_daneshjoo
|
||||||
|
}
|
||||||
|
),
|
||||||
|
style = MaterialTheme.typography.labelSmall,
|
||||||
|
color = MaterialTheme.colorScheme.onSurfaceVariant,
|
||||||
|
)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@ -61,4 +61,8 @@
|
|||||||
<string name="theme_black">سیاه (OLED)</string>
|
<string name="theme_black">سیاه (OLED)</string>
|
||||||
<string name="about">درباره و پروانهها</string>
|
<string name="about">درباره و پروانهها</string>
|
||||||
<string name="about_intro">گنجور برای اندروید نرمافزار آزاد است. شعرها، قلمها و همهٔ کتابخانههایی که این برنامه بر آنها ساخته شده در زیر آمدهاند؛ برای خواندن متن کامل پروانه روی هر مورد بزنید.</string>
|
<string name="about_intro">گنجور برای اندروید نرمافزار آزاد است. شعرها، قلمها و همهٔ کتابخانههایی که این برنامه بر آنها ساخته شده در زیر آمدهاند؛ برای خواندن متن کامل پروانه روی هر مورد بزنید.</string>
|
||||||
|
<string name="no_definition">برای این واژه مدخلی یافت نشد.</string>
|
||||||
|
<string name="source_wiktionary">ویکیواژه (CC BY-SA 3.0)</string>
|
||||||
|
<string name="source_daneshjoo">فرهنگ دانشجو</string>
|
||||||
|
<string name="dictionary">لغتنامه</string>
|
||||||
</resources>
|
</resources>
|
||||||
|
|||||||
@ -61,4 +61,8 @@
|
|||||||
<string name="theme_black">سیاہ (OLED)</string>
|
<string name="theme_black">سیاہ (OLED)</string>
|
||||||
<string name="about">تعارف اور لائسنس</string>
|
<string name="about">تعارف اور لائسنس</string>
|
||||||
<string name="about_intro">گنجور فار اینڈرائیڈ آزاد سافٹ ویئر ہے۔ کلام، فونٹس اور تمام لائبریریاں ذیل میں درج ہیں؛ مکمل لائسنس پڑھنے کے لیے کسی اندراج پر ٹیپ کریں۔</string>
|
<string name="about_intro">گنجور فار اینڈرائیڈ آزاد سافٹ ویئر ہے۔ کلام، فونٹس اور تمام لائبریریاں ذیل میں درج ہیں؛ مکمل لائسنس پڑھنے کے لیے کسی اندراج پر ٹیپ کریں۔</string>
|
||||||
|
<string name="no_definition">اس لفظ کا کوئی اندراج نہیں ملا۔</string>
|
||||||
|
<string name="source_wiktionary">ویکی لغت (CC BY-SA 3.0)</string>
|
||||||
|
<string name="source_daneshjoo">دانشجو لغت</string>
|
||||||
|
<string name="dictionary">لغت نامہ</string>
|
||||||
</resources>
|
</resources>
|
||||||
|
|||||||
@ -66,4 +66,8 @@
|
|||||||
<string name="theme_black">Black (OLED)</string>
|
<string name="theme_black">Black (OLED)</string>
|
||||||
<string name="about">About & licences</string>
|
<string name="about">About & licences</string>
|
||||||
<string name="about_intro">Ganjoor for Android is free software. The poems, the fonts and every library it is built on are credited below; tap an entry to read its full licence.</string>
|
<string name="about_intro">Ganjoor for Android is free software. The poems, the fonts and every library it is built on are credited below; tap an entry to read its full licence.</string>
|
||||||
|
<string name="no_definition">No entry for this word.</string>
|
||||||
|
<string name="source_wiktionary">Wiktionary (CC BY-SA 3.0)</string>
|
||||||
|
<string name="source_daneshjoo">Daneshjoo Dictionary</string>
|
||||||
|
<string name="dictionary">Dictionary</string>
|
||||||
</resources>
|
</resources>
|
||||||
|
|||||||
81
app/src/test/java/com/ganjoor/android/DictionaryTest.kt
Normal file
81
app/src/test/java/com/ganjoor/android/DictionaryTest.kt
Normal file
@ -0,0 +1,81 @@
|
|||||||
|
package com.ganjoor.android
|
||||||
|
|
||||||
|
import com.ganjoor.android.data.affixes
|
||||||
|
import com.ganjoor.android.data.normalise
|
||||||
|
import com.ganjoor.android.data.wordAt
|
||||||
|
import org.junit.Assert.assertEquals
|
||||||
|
import org.junit.Assert.assertNull
|
||||||
|
import org.junit.Assert.assertTrue
|
||||||
|
import org.junit.Test
|
||||||
|
|
||||||
|
class NormaliseTest {
|
||||||
|
@Test
|
||||||
|
fun `diacritics are dropped, since poems carry them and headwords don't`() {
|
||||||
|
assertEquals("الا", normalise("اَلا"))
|
||||||
|
assertEquals("ساقی", normalise("سّاقی"))
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
fun `arabic letter shapes fold to their persian equivalents`() {
|
||||||
|
assertEquals("کی", normalise("كي"))
|
||||||
|
assertEquals("مکه", normalise("مكة"))
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
fun `alef with madda survives, so آتش stays findable`() {
|
||||||
|
assertEquals("آتش", normalise("آتش"))
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
fun `zero-width non-joiners go unless asked for`() {
|
||||||
|
assertEquals("بیروزی", normalise("بیروزی"))
|
||||||
|
assertEquals("بیروزی", normalise("بیروزی", keepZwnj = true))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
class AffixTest {
|
||||||
|
@Test
|
||||||
|
fun `one affix is stripped at a time`() {
|
||||||
|
assertTrue("دلها" .let { affixes(it) }.contains("دل"))
|
||||||
|
assertTrue(affixes("میرود").contains("رود"))
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
fun `a word that is barely longer than its affix is left alone`() {
|
||||||
|
assertEquals(emptyList<String>(), affixes("ها"))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
class WordAtTest {
|
||||||
|
private val line = "اگر آن ترک شیرازی به دست آرد دل ما را"
|
||||||
|
|
||||||
|
@Test
|
||||||
|
fun `a tap inside a word returns that whole word`() {
|
||||||
|
assertEquals("شیرازی", wordAt(line, line.indexOf("شیرازی") + 2))
|
||||||
|
assertEquals("اگر", wordAt(line, 0))
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
fun `a tap on a space returns nothing, so the caller can do something else`() {
|
||||||
|
assertNull(wordAt(line, line.indexOf(' ')))
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
fun `a compound joined by a zero-width non-joiner counts as one word`() {
|
||||||
|
assertEquals("بیخبر", wordAt("سالک بیخبر نبود", 8))
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
fun `single letters and empty text are rejected rather than looked up`() {
|
||||||
|
assertNull(wordAt("", 0))
|
||||||
|
assertNull(wordAt("و دل", 0))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
class ArabicArticleTest {
|
||||||
|
@Test
|
||||||
|
fun `the arabic definite article is stripped, since poems quote arabic`() {
|
||||||
|
assertTrue(affixes("الساقی").contains("ساقی"))
|
||||||
|
assertTrue(affixes("الناس").contains("ناس"))
|
||||||
|
}
|
||||||
|
}
|
||||||
40
tools/README.md
Normal file
40
tools/README.md
Normal file
@ -0,0 +1,40 @@
|
|||||||
|
# Building the dictionary
|
||||||
|
|
||||||
|
`app/src/main/assets/dictionary.db.gz` is generated, not hand-written. This rebuilds it:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
pip install readmdict python-lzo # lzo needs the C library: brew install lzo
|
||||||
|
curl -L -o fa.jsonl https://kaikki.org/dictionary/Persian/kaikki.org-dictionary-Persian.jsonl
|
||||||
|
curl -L -o daneshjoo.mdx \
|
||||||
|
"https://raw.githubusercontent.com/0xdolan/Daneshjoo/main/Daneshjoo%20Dictionary/Daneshjoo%20Dictionary.mdx"
|
||||||
|
python3 build_dictionary.py
|
||||||
|
gzip -9 ganjoor-dictionary.db
|
||||||
|
```
|
||||||
|
|
||||||
|
## Why two sources
|
||||||
|
|
||||||
|
Neither alone is enough. Measured against every distinct word in five real poems (Hafez ×2,
|
||||||
|
Saadi's Golestan, Rumi's Masnavi, a Khayyam rubaʿi — 485 words):
|
||||||
|
|
||||||
|
| | coverage |
|
||||||
|
|---|---|
|
||||||
|
| Daneshjoo alone | 71% |
|
||||||
|
| Wiktionary alone + its form index | ~80% |
|
||||||
|
| **Both, with the lookup chain** | **88%** |
|
||||||
|
|
||||||
|
13 of the 55 remaining misses are Arabic lines quoted inside Persian poems, so Persian coverage
|
||||||
|
is about 91%.
|
||||||
|
|
||||||
|
Wiktionary's export is 93 MB, of which the definitions are 0.94 MB — the rest is inflection
|
||||||
|
tables, etymology templates, IPA and descendants. The build keeps the definitions and the
|
||||||
|
form→lemma index (149,589 pairs) and drops the rest. That index is what resolves conjugated
|
||||||
|
verbs: `افتاد → افتادن`, `بگشاید → گشودن`, `دانند → دانستن`.
|
||||||
|
|
||||||
|
## Licences
|
||||||
|
|
||||||
|
Each row carries its `source`, so attribution stays accurate and either source can be dropped
|
||||||
|
without rebuilding the other.
|
||||||
|
|
||||||
|
- **Wiktionary** — CC BY-SA 3.0. The generated database is therefore also CC BY-SA 3.0.
|
||||||
|
- **Daneshjoo** — the repository states MIT. Note the underlying lexicon is a published Iranian
|
||||||
|
dictionary, so that relicensing is worth verifying before relying on it.
|
||||||
70
tools/build_dictionary.py
Normal file
70
tools/build_dictionary.py
Normal file
@ -0,0 +1,70 @@
|
|||||||
|
"""Builds the app's dictionary from Wiktionary (CC BY-SA 3.0) and Daneshjoo (MIT).
|
||||||
|
|
||||||
|
Wiktionary's export is 93 MB of linguistic metadata around 0.94 MB of definitions, so this keeps
|
||||||
|
the definitions and the form->lemma index and discards the rest. The `source` column is what
|
||||||
|
keeps the attribution honest and lets either source be dropped later.
|
||||||
|
"""
|
||||||
|
import json, re, sqlite3, unicodedata, os, html
|
||||||
|
from readmdict import MDX
|
||||||
|
|
||||||
|
HARAKAT = set(range(0x064B, 0x0653)) | {0x0670, 0x0640} | set(range(0x0610, 0x0616))
|
||||||
|
# Arabic letters that Persian writes differently; headwords and poems disagree constantly.
|
||||||
|
FOLD = {'ي': 'ی', 'ى': 'ی', 'ك': 'ک', 'ة': 'ه'}
|
||||||
|
|
||||||
|
def normalise(s: str) -> str:
|
||||||
|
d = unicodedata.normalize('NFD', s)
|
||||||
|
d = ''.join(c for c in d if ord(c) not in HARAKAT)
|
||||||
|
s = unicodedata.normalize('NFC', d)
|
||||||
|
return ''.join(FOLD.get(c, c) for c in s).replace('', '').strip()
|
||||||
|
|
||||||
|
db = 'ganjoor-dictionary.db'
|
||||||
|
if os.path.exists(db): os.remove(db)
|
||||||
|
c = sqlite3.connect(db)
|
||||||
|
c.executescript("""
|
||||||
|
CREATE TABLE entry (word TEXT NOT NULL, display TEXT NOT NULL, gloss TEXT NOT NULL, source TEXT NOT NULL);
|
||||||
|
CREATE TABLE form (form TEXT NOT NULL, lemma TEXT NOT NULL);
|
||||||
|
""")
|
||||||
|
|
||||||
|
entries, forms = [], set()
|
||||||
|
|
||||||
|
for line in open('fa.jsonl', encoding='utf-8'):
|
||||||
|
try: e = json.loads(line)
|
||||||
|
except Exception: continue
|
||||||
|
word = e.get('word')
|
||||||
|
if not word: continue
|
||||||
|
gs = [g.strip() for s in e.get('senses', []) for g in (s.get('glosses') or []) if g.strip()]
|
||||||
|
if gs:
|
||||||
|
pos = e.get('pos') or ''
|
||||||
|
gloss = '; '.join(dict.fromkeys(gs))[:600]
|
||||||
|
entries.append((normalise(word), word, f"({pos}) {gloss}" if pos else gloss, 'wiktionary'))
|
||||||
|
for f in e.get('forms', []):
|
||||||
|
t = f.get('form')
|
||||||
|
if t and t != word and not t.startswith('-') and len(t) > 1:
|
||||||
|
forms.add((normalise(t), normalise(word)))
|
||||||
|
|
||||||
|
print(f"wiktionary: {len(entries)} entries, {len(forms)} forms")
|
||||||
|
|
||||||
|
tag = re.compile(r'<[^>]+>')
|
||||||
|
n0 = len(entries)
|
||||||
|
for k, v in MDX('daneshjoo.mdx').items():
|
||||||
|
word = k.decode('utf-8', 'ignore')
|
||||||
|
if not re.match(r'^[-ۿ]', word):
|
||||||
|
continue # the en->fa half isn't useful here
|
||||||
|
txt = html.unescape(tag.sub(' ', v.decode('utf-8', 'ignore')))
|
||||||
|
txt = re.sub(r'\s+', ' ', txt).strip()
|
||||||
|
if txt.startswith(word):
|
||||||
|
txt = txt[len(word):].strip(' -–—')
|
||||||
|
if txt:
|
||||||
|
entries.append((normalise(word), word, txt[:600], 'daneshjoo'))
|
||||||
|
print(f"daneshjoo : {len(entries) - n0} entries")
|
||||||
|
|
||||||
|
c.executemany("INSERT INTO entry VALUES (?,?,?,?)", entries)
|
||||||
|
c.executemany("INSERT INTO form VALUES (?,?)", sorted(forms))
|
||||||
|
c.executescript("""
|
||||||
|
CREATE INDEX idx_entry_word ON entry(word);
|
||||||
|
CREATE INDEX idx_form_form ON form(form);
|
||||||
|
""")
|
||||||
|
c.commit()
|
||||||
|
c.execute("VACUUM")
|
||||||
|
c.close()
|
||||||
|
print(f"\n{db}: {os.path.getsize(db)/1e6:.1f} MB")
|
||||||
Loading…
Reference in New Issue
Block a user