PWG literary-source scan-index campaign
Between January 2025 and July 2026 a volunteer team page-indexed the printed
editions that the Böhtlingk-Roth Sanskrit-Wörterbuch (PWG) cites, so that an
<ls> citation in the dictionary can be resolved to the page image of the edition
it cites. This page measures that campaign by citation mass — a 159-page
Kumārasaṃbhava and a 2,420-page Taittirīyabrāhmaṇa are neither equal work nor
equal payoff.
Complement, not duplicate, of Citation Coverage: that page asks how many citations link out; this one asks how the link targets got built.
Works indexed
Citation mass indexed
Pages indexed
Volunteers
:::note
Scope.
Where the work stands
Status is the sheet's own vocabulary. Two of its values are rulings, not backlog:
page-wise marks a work PWG cites by page rather than by verse, so a per-entry index
would answer a question nobody asks of it; nr-* marks an abbreviation that is an
indirect citation or an alternate name for a work already indexed.
How to read: Each bar is one status; length is the citation mass of the works in it, so the chart shows payoff at stake, not headcount. Example 1:
donedwarfing every other bar means the campaign has already captured most of the citation value it set out to capture. Example 2:page-wisebeing large is not a backlog warning — it is the mass the campaign deliberately declined to index.
Conclusion: The campaign finished the works that carry the citation weight. What is left unclaimed is a small share of the tracked mass, concentrated in long Vedic texts.
Velocity — two pipelines, not one
An index being finished by its volunteer and its scan directory going public are separate steps with separate queues. Plotting them together shows the publishing lag as a visible offset rather than hiding it inside one "progress" line.
How to read: Two series over the campaign's months — indexes finished, and scan directories published. Example 1: A tall finished-bar with no published-bar beside it is a month whose output was still in the publishing queue. Example 2: Published exceeding finished in a later month is that queue draining.
The median lag from an index being posted to its scan directory going public is
Conclusion: Throughput peaked in February–March 2025 and decayed through the year — a volunteer campaign's normal shape, not a stall. The publishing pipeline tracked the indexing pipeline with a lag of under two weeks at the median.
Who did it
Attribution is exactly the sheet's Reserved/Indexed by column. Rows count assigned
work, so a volunteer's row count includes work in progress, not only finished indexes.
How to read: Each bar is one volunteer; length is the citation mass of the works they took. Example 1: Two volunteers carrying a third of the mass between them is the bus-factor pattern this observatory measures elsewhere, reproduced inside a single campaign. Example 2: A volunteer with many pages but little citation mass took long, thinly-cited books — real work the mass metric under-credits.
Conclusion: Eight volunteers carried the campaign, with the top three taking roughly half the citation mass between them. One volunteer's work sits entirely in multi-volume books whose citation count the sheet records only once, on volume 1 — their mass reads as zero here and their page count tells the truer story.
What remains
The unclaimed backlog, ranked by citation payoff. The ★ marking in the source sheet picks out exactly the Vedic saṃhitā / brāhmaṇa / upaniṣad / śrauta- and gṛhya-sūtra / prātiśākhya group.
How to read: One bar per unclaimed work, longest citation count first. Example 1: A short bar over a large page count is a low-yield, high-effort target. Example 2: The colour split shows how much of what is left is Vedic — the material with the awkward reference schemes.
Conclusion: The remaining work is Vedic and long. The kāvya and kośa material, which indexes quickly and is cited densely, is finished — so the backlog's citation payoff per page is the lowest the campaign has faced.
Source
- report —
reports/pwg_scan_index.md - registry —
data/pwg_scan_index_tracker/pwg_scan_index.tsv - e-text candidate queue —
pwg_etext_candidate_queue.tsv - campaign history —
docs/PWG_SCAN_INDEX_CAMPAIGN_HISTORY_2025_2026.md - generator —
scripts/pwg_scan_index.py - upstream tracker — Google Sheet, snapshotted verbatim under
snapshot/