Go to file
Anas Rashid 256800c577
Some checks are pending
Wikisource sync / sync (push) Waiting to run
Published order for books and poems; divan (radif letter) order for ghazals
- wikisource.py: record each work's position in the author page list (ord), refreshed on
  every sync; children of index pages keep their parent's position
- export: walk works by ord; books in order of first appearance (Iqbal: Bang-e-Dra,
  Bal-e-Jibril, Zarb-e-Kalim, Armaghan-e-Hijaz)
- ghazal sections in divan order: last letter of the radif (opening line), alif..ye;
  ties keep source order; Iqbal keeps his books' published order

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-08 20:45:45 +02:00
.gitea/workflows Rename export_ganjoor.py -> export_divan.py, ganjoor_ids -> divan_ids; Language ur-PK 2026-10-04 23:51:03 +02:00
export Published order for books and poems; divan (radif letter) order for ghazals 2026-10-08 20:45:45 +02:00
index Merge duplicate author pages via resolved Wikipedia article (#7) 2026-10-05 00:32:32 +02:00
poets Published order for books and poems; divan (radif letter) order for ghazals 2026-10-08 20:45:45 +02:00
.gitignore Divan: classical Urdu poetry & prose from Urdu Wikisource 2026-10-04 22:48:40 +02:00
build_index.py Urdu only: drop English Wikipedia intros (intro_ur -> intro) 2026-10-04 22:52:15 +02:00
divan.db Published order for books and poems; divan (radif letter) order for ghazals 2026-10-08 20:45:45 +02:00
export_divan.py Published order for books and poems; divan (radif letter) order for ghazals 2026-10-08 20:45:45 +02:00
languages.json Ganjoor-compatible export (manifest.json, poets/, index/) 2026-10-04 23:05:14 +02:00
LICENSE Divan: classical Urdu poetry & prose from Urdu Wikisource 2026-10-04 22:48:40 +02:00
manifest.json Published order for books and poems; divan (radif letter) order for ghazals 2026-10-08 20:45:45 +02:00
metres.json Ganjoor-compatible export (manifest.json, poets/, index/) 2026-10-04 23:05:14 +02:00
README.md Rename export_ganjoor.py -> export_divan.py, ganjoor_ids -> divan_ids; Language ur-PK 2026-10-04 23:51:03 +02:00
update.sh Rename export_ganjoor.py -> export_divan.py, ganjoor_ids -> divan_ids; Language ur-PK 2026-10-04 23:51:03 +02:00
wikisource.py Published order for books and poems; divan (radif letter) order for ghazals 2026-10-08 20:45:45 +02:00

دیوان · Divan

A local, searchable SQLite database of classical Urdu literature, both poetry and prose: ghazals, nazms, marsiyas, masnavis, letters, dastans and essays by poets and writers such as Mir, Sauda, Dard, Ghalib, Momin, Zauq, Dagh, Anees, Hali, Akbar Allahabadi, Allama Iqbal, Mir Amman, Sir Syed and Nazir Ahmad. Each author entry includes a short introduction.

Everything is in Urdu script.

Sources

  • Texts: Urdu Wikisource (ویکی ماخذ). The works are in the public domain ({{PD-old}}, {{PD-Pakistan}}, …), and the license tag is kept on every record.
  • Author intros: the lead section of each author's Urdu Wikipedia article.

Everything is fetched through the official MediaWiki API.

Database (divan.db)

table contents
poets page, name, years, birth_year, death_year, description, image, wikipedia, wikidata, intro, url
works title, poet_page, kind (poetry / prose), section (genre or book, e.g. شاعری > بانگ درا (1924)), year, text_ur, license, url
works_fts, poets_fts SQLite FTS5 full-text indexes

In poetry, lines are separated by \n and couplets or stanzas by a blank line.

Usage

python3 wikisource.py    # first build ~30 min; later runs fetch only pages edited since the last run
python3 build_index.py   # full-text index + export/poets.jsonl, export/works.jsonl

Python 3.9+ standard library only. (Search needs FTS5. Python's sqlite3 has it; Apple's built-in sqlite3 command does not, so use Python or brew install sqlite.)

A Gitea Actions workflow (.gitea/workflows/sync.yml) runs update.sh daily. It commits only when the data changed.

-- Iqbal's Bang-e-Dra
SELECT title, text_ur FROM works WHERE poet_page='مصنف:محمد اقبال' AND section LIKE '%بانگ درا%';

-- all prose
SELECT poet_page, title FROM works WHERE kind='prose';

-- full-text search
SELECT title, poet_page FROM works_fts WHERE works_fts MATCH 'خودی';

Site/API format (Ganjoor-compatible)

The repo root also holds the data in the ganjoor-data layout (manifest.json, poets/, index/, built by export_divan.py). It works as a static API over jsDelivr with no server:

https://cdn.jsdelivr.net/gh/anas-rashid/divan-data@main/manifest.json
https://cdn.jsdelivr.net/gh/anas-rashid/divan-data@main/poets/p238/_cat.json        # Iqbal

It also loads directly into GanjoorService through its "public data import" page: give it the base URL above. Poets are /p{id}, categories are a genre slug (ghazal, nazm, …) or c{id}, and poems are sh{id}. Ids stay stable across syncs. Poet years are converted to approximate Hijri, following Ganjoor's convention.

License

  • Code: MIT.
  • Data (divan.db, export/): the literary works are in the public domain. The compilation is derived from Wikisource and Wikipedia and is shared under CC BY-SA 4.0. When reusing it, credit "Urdu Wikisource and Wikipedia contributors" and share alike. Every record keeps its source url.

Corrections belong upstream: fix the text on Wikisource and re-run the script.