Divan: classical Urdu poetry & prose from Urdu Wikisource
384 authors, 11,087 works (10,446 poetry, 641 prose), FTS5 search, JSONL export, daily Gitea Actions sync. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
commit
81ab6a3331
14
.gitea/workflows/sync.yml
Normal file
14
.gitea/workflows/sync.yml
Normal file
@ -0,0 +1,14 @@
|
|||||||
|
name: Wikisource sync
|
||||||
|
on:
|
||||||
|
schedule:
|
||||||
|
- cron: "0 3 * * *" # daily 03:00 UTC
|
||||||
|
workflow_dispatch:
|
||||||
|
jobs:
|
||||||
|
sync:
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@v4
|
||||||
|
- run: |
|
||||||
|
git config user.name "divan-bot"
|
||||||
|
git config user.email "divan-bot@noreply.git.anasrashid.net"
|
||||||
|
./update.sh
|
||||||
4
.gitignore
vendored
Normal file
4
.gitignore
vendored
Normal file
@ -0,0 +1,4 @@
|
|||||||
|
__pycache__/
|
||||||
|
*.db-journal
|
||||||
|
wikisource.log
|
||||||
|
.DS_Store
|
||||||
21
LICENSE
Normal file
21
LICENSE
Normal file
@ -0,0 +1,21 @@
|
|||||||
|
MIT License
|
||||||
|
|
||||||
|
Copyright (c) 2026 Anas Rashid
|
||||||
|
|
||||||
|
Permission is hereby granted, free of charge, to any person obtaining a copy
|
||||||
|
of this software and associated documentation files (the "Software"), to deal
|
||||||
|
in the Software without restriction, including without limitation the rights
|
||||||
|
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
||||||
|
copies of the Software, and to permit persons to whom the Software is
|
||||||
|
furnished to do so, subject to the following conditions:
|
||||||
|
|
||||||
|
The above copyright notice and this permission notice shall be included in all
|
||||||
|
copies or substantial portions of the Software.
|
||||||
|
|
||||||
|
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||||
|
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
||||||
|
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
||||||
|
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
||||||
|
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
||||||
|
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
||||||
|
SOFTWARE.
|
||||||
51
README.md
Normal file
51
README.md
Normal file
@ -0,0 +1,51 @@
|
|||||||
|
# دیوان · Divan
|
||||||
|
|
||||||
|
A local, searchable SQLite database of **classical Urdu literature, both poetry and prose**: ghazals, nazms, marsiyas, masnavis, letters, dastans and essays by poets and writers such as Mir, Sauda, Dard, Ghalib, Momin, Zauq, Dagh, Anees, Hali, Akbar Allahabadi, Allama Iqbal, Mir Amman, Sir Syed and Nazir Ahmad. Each author entry includes a short introduction.
|
||||||
|
|
||||||
|
Texts are in **Urdu script**. Author intros are in Urdu, plus English where available.
|
||||||
|
|
||||||
|
## Sources
|
||||||
|
|
||||||
|
- **Texts:** [Urdu Wikisource](https://ur.wikisource.org) (ویکی ماخذ). The works are in the public domain (`{{PD-old}}`, `{{PD-Pakistan}}`, …), and the license tag is kept on every record.
|
||||||
|
- **Author intros:** the lead section of each author's [Urdu Wikipedia](https://ur.wikipedia.org) and [English Wikipedia](https://en.wikipedia.org) article.
|
||||||
|
|
||||||
|
Everything is fetched through the official MediaWiki API.
|
||||||
|
|
||||||
|
## Database (`divan.db`)
|
||||||
|
|
||||||
|
| table | contents |
|
||||||
|
|---|---|
|
||||||
|
| `poets` | `page`, `name`, `years`, `birth_year`, `death_year`, `description`, `image`, `wikipedia`, `wikidata`, `intro_ur`, `intro_en`, `url` |
|
||||||
|
| `works` | `title`, `poet_page`, `kind` (`poetry` / `prose`), `section` (genre or book, e.g. `شاعری > بانگ درا (1924)`), `year`, `text_ur`, `license`, `url` |
|
||||||
|
| `works_fts`, `poets_fts` | SQLite FTS5 full-text indexes |
|
||||||
|
|
||||||
|
In poetry, lines are separated by `\n` and couplets or stanzas by a blank line.
|
||||||
|
|
||||||
|
## Usage
|
||||||
|
|
||||||
|
```sh
|
||||||
|
python3 wikisource.py # first build ~30 min; later runs fetch only pages edited since the last run
|
||||||
|
python3 build_index.py # full-text index + export/poets.jsonl, export/works.jsonl
|
||||||
|
```
|
||||||
|
|
||||||
|
Python 3.9+ standard library only. (Search needs FTS5. Python's `sqlite3` has it; Apple's built-in `sqlite3` command does not, so use Python or `brew install sqlite`.)
|
||||||
|
|
||||||
|
A Gitea Actions workflow (`.gitea/workflows/sync.yml`) runs `update.sh` daily. It commits only when the data changed.
|
||||||
|
|
||||||
|
```sql
|
||||||
|
-- Iqbal's Bang-e-Dra
|
||||||
|
SELECT title, text_ur FROM works WHERE poet_page='مصنف:محمد اقبال' AND section LIKE '%بانگ درا%';
|
||||||
|
|
||||||
|
-- all prose
|
||||||
|
SELECT poet_page, title FROM works WHERE kind='prose';
|
||||||
|
|
||||||
|
-- full-text search
|
||||||
|
SELECT title, poet_page FROM works_fts WHERE works_fts MATCH 'خودی';
|
||||||
|
```
|
||||||
|
|
||||||
|
## License
|
||||||
|
|
||||||
|
- **Code:** [MIT](LICENSE).
|
||||||
|
- **Data** (`divan.db`, `export/`): the literary works are in the public domain. The compilation is derived from Wikisource and Wikipedia and is shared under **[CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/)**. When reusing it, credit *"Urdu Wikisource and Wikipedia contributors"* and share alike. Every record keeps its source `url`.
|
||||||
|
|
||||||
|
Corrections belong upstream: fix the text on Wikisource and re-run the script.
|
||||||
26
build_index.py
Normal file
26
build_index.py
Normal file
@ -0,0 +1,26 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Build FTS5 full-text search over divan.db and export JSONL. Safe to re-run."""
|
||||||
|
import json, os, sqlite3
|
||||||
|
|
||||||
|
d = os.path.dirname(os.path.abspath(__file__))
|
||||||
|
db = sqlite3.connect(f"{d}/divan.db")
|
||||||
|
db.executescript("""
|
||||||
|
DROP TABLE IF EXISTS works_fts;
|
||||||
|
CREATE VIRTUAL TABLE works_fts USING fts5(title, poet_page UNINDEXED, kind UNINDEXED, section, text_ur,
|
||||||
|
content='works', tokenize='unicode61 remove_diacritics 2');
|
||||||
|
INSERT INTO works_fts(rowid, title, poet_page, kind, section, text_ur) SELECT rowid, title, poet_page, kind, section, text_ur FROM works;
|
||||||
|
DROP TABLE IF EXISTS poets_fts;
|
||||||
|
CREATE VIRTUAL TABLE poets_fts USING fts5(page UNINDEXED, name, description, intro_ur, intro_en,
|
||||||
|
content='poets', tokenize='unicode61 remove_diacritics 2');
|
||||||
|
INSERT INTO poets_fts(rowid, page, name, description, intro_ur, intro_en) SELECT rowid, page, name, description, intro_ur, intro_en FROM poets;
|
||||||
|
""")
|
||||||
|
db.commit()
|
||||||
|
db.execute("VACUUM")
|
||||||
|
|
||||||
|
os.makedirs(f"{d}/export", exist_ok=True)
|
||||||
|
db.row_factory = sqlite3.Row
|
||||||
|
for t in ("poets", "works"):
|
||||||
|
with open(f"{d}/export/{t}.jsonl", "w", encoding="utf-8") as f:
|
||||||
|
for r in db.execute(f"SELECT * FROM {t} ORDER BY 1"):
|
||||||
|
f.write(json.dumps(dict(r), ensure_ascii=False) + "\n")
|
||||||
|
print({t: db.execute(f"SELECT count(*) FROM {t}").fetchone()[0] for t in ("poets", "works")})
|
||||||
384
export/poets.jsonl
Normal file
384
export/poets.jsonl
Normal file
File diff suppressed because one or more lines are too long
11087
export/works.jsonl
Normal file
11087
export/works.jsonl
Normal file
File diff suppressed because one or more lines are too long
12
update.sh
Executable file
12
update.sh
Executable file
@ -0,0 +1,12 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# Pull new/edited pages from Wikisource, rebuild index, commit + push only if the data changed.
|
||||||
|
set -euo pipefail
|
||||||
|
cd "$(dirname "$0")"
|
||||||
|
python3 wikisource.py
|
||||||
|
python3 build_index.py
|
||||||
|
git add export
|
||||||
|
if git diff --cached --quiet -- export; then echo "no changes"; exit 0; fi # export/*.jsonl is the change signal; divan.db bytes change every run
|
||||||
|
git add divan.db
|
||||||
|
git commit -q -m "data: Wikisource sync $(date -u +%F)"
|
||||||
|
git push -q origin HEAD
|
||||||
|
echo "pushed"
|
||||||
208
wikisource.py
Normal file
208
wikisource.py
Normal file
@ -0,0 +1,208 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Build divan.db from Urdu Wikisource (public-domain texts, CC BY-SA 4.0 site)
|
||||||
|
plus short poet intros from Urdu/English Wikipedia. Uses the MediaWiki API, 50 pages per request."""
|
||||||
|
import json, os, re, sqlite3, time, urllib.parse, urllib.request
|
||||||
|
|
||||||
|
WS = "https://ur.wikisource.org/w/api.php"
|
||||||
|
DB = os.path.join(os.path.dirname(os.path.abspath(__file__)), "divan.db")
|
||||||
|
UA = "divan-dataset/0.1 (non-commercial Urdu poetry archive; python-urllib)"
|
||||||
|
|
||||||
|
db = sqlite3.connect(DB)
|
||||||
|
db.executescript("""
|
||||||
|
CREATE TABLE IF NOT EXISTS poets(page TEXT PRIMARY KEY, name TEXT, years TEXT, birth_year TEXT, death_year TEXT,
|
||||||
|
description TEXT, image TEXT, wikipedia TEXT, wikidata TEXT, intro_ur TEXT, intro_en TEXT, url TEXT);
|
||||||
|
CREATE TABLE IF NOT EXISTS works(title TEXT PRIMARY KEY, poet_page TEXT, kind TEXT, section TEXT, year TEXT,
|
||||||
|
text_ur TEXT, license TEXT, url TEXT);
|
||||||
|
CREATE INDEX IF NOT EXISTS works_poet ON works(poet_page);
|
||||||
|
CREATE TABLE IF NOT EXISTS meta(key TEXT PRIMARY KEY, value TEXT);
|
||||||
|
""")
|
||||||
|
|
||||||
|
|
||||||
|
def api(base, **params):
|
||||||
|
params.update(format="json", formatversion=2, maxlag=5)
|
||||||
|
for i in range(5):
|
||||||
|
try:
|
||||||
|
req = urllib.request.Request(base, data=urllib.parse.urlencode(params).encode(), headers={"User-Agent": UA}) # POST: 50 Urdu titles overflow a GET URL
|
||||||
|
with urllib.request.urlopen(req, timeout=60) as r:
|
||||||
|
d = json.load(r)
|
||||||
|
if d.get("error", {}).get("code") == "maxlag":
|
||||||
|
raise RuntimeError("maxlag")
|
||||||
|
return d
|
||||||
|
except Exception:
|
||||||
|
time.sleep(5 * (i + 1))
|
||||||
|
raise RuntimeError(f"API failed: {params}")
|
||||||
|
|
||||||
|
|
||||||
|
def wikitexts(titles):
|
||||||
|
"""{requested title: (resolved title, wikitext)} following redirects, 50 per request."""
|
||||||
|
out = {}
|
||||||
|
for i in range(0, len(titles), 50):
|
||||||
|
chunk = titles[i:i + 50]
|
||||||
|
q = api(WS, action="query", prop="revisions", rvprop="content", rvslots="main",
|
||||||
|
titles="|".join(chunk), redirects=1)["query"]
|
||||||
|
alias = {}
|
||||||
|
for k in ("normalized", "redirects"):
|
||||||
|
for r in q.get(k, []):
|
||||||
|
alias[r["from"]] = r["to"]
|
||||||
|
pages = {p["title"]: p["revisions"][0]["slots"]["main"]["content"]
|
||||||
|
for p in q.get("pages", []) if "revisions" in p}
|
||||||
|
for t in chunk:
|
||||||
|
r = t
|
||||||
|
while r in alias:
|
||||||
|
r = alias[r]
|
||||||
|
if r in pages:
|
||||||
|
out[t] = (r, pages[r])
|
||||||
|
time.sleep(0.5)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def tpl_field(text, name):
|
||||||
|
m = re.search(rf"^\s*\|\s*{name}\s*=(.*)$", text, re.M)
|
||||||
|
return m.group(1).strip() or None if m else None
|
||||||
|
|
||||||
|
|
||||||
|
def page_url(t):
|
||||||
|
return "https://ur.wikisource.org/wiki/" + urllib.parse.quote(t.replace(" ", "_"))
|
||||||
|
|
||||||
|
|
||||||
|
def clean(s):
|
||||||
|
s = re.sub(r"<!--.*?-->|<ref[^>]*>.*?</ref>|<ref[^>]*/>", "", s, flags=re.S)
|
||||||
|
s = re.sub(r"\{\{[^{}]*\}\}", "", s)
|
||||||
|
s = re.sub(r"\[\[(?:[^|\]]*\|)?([^\]]*)\]\]", r"\1", s)
|
||||||
|
s = re.sub(r"'''?|<[^>]+>||", "", s)
|
||||||
|
return re.sub(r"\n{3,}", "\n\n", "\n".join(l.rstrip() for l in s.splitlines())).strip()
|
||||||
|
|
||||||
|
|
||||||
|
def poem_text(wt):
|
||||||
|
poems = re.findall(r"<poem[^>]*>(.*?)</poem>", wt, re.S)
|
||||||
|
if poems:
|
||||||
|
return "\n\n".join(clean(p) for p in poems)
|
||||||
|
body = re.sub(r"\{\{\s*header.*?\n\}\}", "", wt, flags=re.S | re.I)
|
||||||
|
body = re.sub(r"\[\[(Category|زمرہ):[^\]]*\]\]", "", body)
|
||||||
|
return clean(body)
|
||||||
|
|
||||||
|
|
||||||
|
def author_links(wt):
|
||||||
|
"""[(target, section path)] from an author page's bullet lists."""
|
||||||
|
path, out = {}, []
|
||||||
|
for line in wt.splitlines():
|
||||||
|
h = re.match(r"^(=+)\s*(.*?)\s*\1\s*$", line)
|
||||||
|
if h:
|
||||||
|
lvl = len(h.group(1))
|
||||||
|
path = {k: v for k, v in path.items() if k < lvl}
|
||||||
|
path[lvl] = clean(h.group(2))
|
||||||
|
continue
|
||||||
|
if line.startswith("*"):
|
||||||
|
for t in re.findall(r"\[\[([^|\]#]+)", line):
|
||||||
|
t = t.strip()
|
||||||
|
if ":" not in t and not t.startswith("/"):
|
||||||
|
out.append((t, " > ".join(v for k, v in sorted(path.items()) if v != "تصانیف")))
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def all_authors():
|
||||||
|
titles, cont = [], {}
|
||||||
|
while True:
|
||||||
|
d = api(WS, action="query", list="allpages", apnamespace=102, aplimit=500, **cont)
|
||||||
|
titles += [p["title"] for p in d["query"]["allpages"]]
|
||||||
|
if "continue" not in d:
|
||||||
|
return titles
|
||||||
|
cont = {"apcontinue": d["continue"]["apcontinue"]}
|
||||||
|
|
||||||
|
|
||||||
|
def intros(wp_titles, lang):
|
||||||
|
"""Plain-text lead section from {lang}.wikipedia, 20 per request."""
|
||||||
|
out = {}
|
||||||
|
base = f"https://{lang}.wikipedia.org/w/api.php"
|
||||||
|
for i in range(0, len(wp_titles), 20):
|
||||||
|
chunk = wp_titles[i:i + 20]
|
||||||
|
q = api(base, action="query", prop="extracts|langlinks", exintro=1, explaintext=1, exlimit=20,
|
||||||
|
lllang="en", titles="|".join(chunk), redirects=1)["query"]
|
||||||
|
alias = {r["from"]: r["to"] for k in ("normalized", "redirects") for r in q.get(k, [])}
|
||||||
|
pages = {p["title"]: p for p in q.get("pages", [])}
|
||||||
|
for t in chunk:
|
||||||
|
r = alias.get(alias.get(t, t), alias.get(t, t))
|
||||||
|
if r in pages:
|
||||||
|
out[t] = pages[r]
|
||||||
|
time.sleep(0.5)
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def changed_since(ts):
|
||||||
|
"""Titles in the main namespace edited/created on Wikisource since ts (recentchanges keeps ~30 days)."""
|
||||||
|
titles, cont = set(), {}
|
||||||
|
while True:
|
||||||
|
d = api(WS, action="query", list="recentchanges", rcnamespace=0, rcdir="newer", rcstart=ts,
|
||||||
|
rcprop="title", rclimit=500, **cont)
|
||||||
|
titles |= {c["title"] for c in d["query"]["recentchanges"]}
|
||||||
|
if "continue" not in d:
|
||||||
|
return titles
|
||||||
|
cont = {"rccontinue": d["continue"]["rccontinue"]}
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
started = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
|
||||||
|
last = (db.execute("SELECT value FROM meta WHERE key='last_run'").fetchone() or [None])[0]
|
||||||
|
authors = all_authors()
|
||||||
|
print(len(authors), "author pages")
|
||||||
|
apages = wikitexts(authors)
|
||||||
|
links = [] # (work title, poet page, section)
|
||||||
|
for a, (_, wt) in apages.items():
|
||||||
|
f = lambda n: tpl_field(wt, n)
|
||||||
|
wp = (f("wikipedia") or "").removeprefix("ur:") or None
|
||||||
|
db.execute("INSERT OR REPLACE INTO poets(page,name,years,birth_year,death_year,description,image,wikipedia,wikidata,url) "
|
||||||
|
"VALUES (?,?,?,?,?,?,?,?,?,?)",
|
||||||
|
(a, f("firstname") or a.split(":", 1)[1], f("dates"), f("birthyear"), f("deathyear"),
|
||||||
|
f("description"), f("image"), wp, f("wikidata"), page_url(a)))
|
||||||
|
links += [(t, a, s) for t, s in author_links(wt)]
|
||||||
|
db.commit()
|
||||||
|
|
||||||
|
# Wikipedia intros (ur, then en via langlinks)
|
||||||
|
wps = [r[0] for r in db.execute("SELECT DISTINCT wikipedia FROM poets WHERE wikipedia IS NOT NULL")]
|
||||||
|
ur = intros(wps, "ur")
|
||||||
|
en_titles = {t: p["langlinks"][0]["title"] for t, p in ur.items() if p.get("langlinks")}
|
||||||
|
en = intros(list(set(en_titles.values())), "en")
|
||||||
|
for t, p in ur.items():
|
||||||
|
e = en.get(en_titles.get(t), {})
|
||||||
|
db.execute("UPDATE poets SET intro_ur=?, intro_en=? WHERE wikipedia=?", (p.get("extract"), e.get("extract"), t))
|
||||||
|
db.commit()
|
||||||
|
print(len(ur), "ur intros,", len(en), "en intros")
|
||||||
|
|
||||||
|
# Works; index-like pages (no <poem>, mostly links) are expanded one level
|
||||||
|
seen = {r[0] for r in db.execute("SELECT title FROM works")}
|
||||||
|
if last: # incremental: refetch pages edited since the last run
|
||||||
|
changed = changed_since(last)
|
||||||
|
seen -= changed
|
||||||
|
print(f"{len(changed)} pages changed since {last}", flush=True)
|
||||||
|
queue, depth = links, 0
|
||||||
|
while queue and depth < 3:
|
||||||
|
todo = {}
|
||||||
|
for t, a, s in queue:
|
||||||
|
todo.setdefault(t, (a, s))
|
||||||
|
pending = [t for t in todo if t not in seen]
|
||||||
|
print(f"depth {depth}: {len(pending)} pages to fetch", flush=True)
|
||||||
|
nxt = []
|
||||||
|
for i in range(0, len(pending), 500): # commit + report every 500 pages
|
||||||
|
texts = wikitexts(pending[i:i + 500])
|
||||||
|
for t, (title, wt) in texts.items():
|
||||||
|
a, s = todo[t]
|
||||||
|
if title in seen:
|
||||||
|
continue
|
||||||
|
seen.add(title)
|
||||||
|
if "<poem" not in wt and len(re.findall(r"^\*\s*\[\[", wt, re.M)) >= 3:
|
||||||
|
nxt += [(c, a, f"{s} > {title}".strip(" >")) for c, _ in author_links(wt)]
|
||||||
|
nxt += [(title + c, a, f"{s} > {title}".strip(" >"))
|
||||||
|
for c in re.findall(r"\[\[(/[^|\]]+)", wt)]
|
||||||
|
continue
|
||||||
|
lic = re.findall(r"\{\{\s*(PD[^}|]*)", wt)
|
||||||
|
db.execute("INSERT OR REPLACE INTO works VALUES (?,?,?,?,?,?,?,?)",
|
||||||
|
(title, a, "poetry" if "<poem" in wt else "prose", s or None, tpl_field(wt, "year"), poem_text(wt), lic[0].strip() if lic else None, page_url(title)))
|
||||||
|
db.commit()
|
||||||
|
print(f" {min(i + 500, len(pending))}/{len(pending)}, works total {db.execute('SELECT count(*) FROM works').fetchone()[0]}", flush=True)
|
||||||
|
queue, depth = nxt, depth + 1
|
||||||
|
db.execute("INSERT OR REPLACE INTO meta VALUES ('last_run', ?)", (started,))
|
||||||
|
db.commit()
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
Loading…
Reference in New Issue
Block a user