CSL Observatory 13 years of Cologne Digital Sanskrit Lexicon

PWG literary-source scan-index campaign

Between January 2025 and July 2026 a volunteer team page-indexed the printed editions that the Böhtlingk-Roth Sanskrit-Wörterbuch (PWG) cites, so that an <ls> citation in the dictionary can be resolved to the page image of the edition it cites. This page measures that campaign by citation mass — a 159-page Kumārasaṃbhava and a 2,420-page Taittirīyabrāhmaṇa are neither equal work nor equal payoff.

Complement, not duplicate, of Citation Coverage: that page asks how many citations link out; this one asks how the link targets got built.

Works indexed

/

Citation mass indexed

%

Pages indexed

Volunteers

:::note Scope. literary sources tracked, snapshot . The percentage is coverage of the tracked set — the works the campaign took on — not of the dictionary's whole citation apparatus. Those two denominators are not interchangeable, and the full report §6.2 shows why. :::

Where the work stands

Status is the sheet's own vocabulary. Two of its values are rulings, not backlog: page-wise marks a work PWG cites by page rather than by verse, so a per-entry index would answer a question nobody asks of it; nr-* marks an abbreviation that is an indirect citation or an alternate name for a work already indexed.

How to read: Each bar is one status; length is the citation mass of the works in it, so the chart shows payoff at stake, not headcount. Example 1: done dwarfing every other bar means the campaign has already captured most of the citation value it set out to capture. Example 2: page-wise being large is not a backlog warning — it is the mass the campaign deliberately declined to index.

Conclusion: The campaign finished the works that carry the citation weight. What is left unclaimed is a small share of the tracked mass, concentrated in long Vedic texts.

Velocity — two pipelines, not one

An index being finished by its volunteer and its scan directory going public are separate steps with separate queues. Plotting them together shows the publishing lag as a visible offset rather than hiding it inside one "progress" line.

How to read: Two series over the campaign's months — indexes finished, and scan directories published. Example 1: A tall finished-bar with no published-bar beside it is a month whose output was still in the publishing queue. Example 2: Published exceeding finished in a later month is that queue draining.

The median lag from an index being posted to its scan directory going public is days.

Conclusion: Throughput peaked in February–March 2025 and decayed through the year — a volunteer campaign's normal shape, not a stall. The publishing pipeline tracked the indexing pipeline with a lag of under two weeks at the median.

Who did it

Attribution is exactly the sheet's Reserved/Indexed by column. Rows count assigned work, so a volunteer's row count includes work in progress, not only finished indexes.

How to read: Each bar is one volunteer; length is the citation mass of the works they took. Example 1: Two volunteers carrying a third of the mass between them is the bus-factor pattern this observatory measures elsewhere, reproduced inside a single campaign. Example 2: A volunteer with many pages but little citation mass took long, thinly-cited books — real work the mass metric under-credits.

Conclusion: Eight volunteers carried the campaign, with the top three taking roughly half the citation mass between them. One volunteer's work sits entirely in multi-volume books whose citation count the sheet records only once, on volume 1 — their mass reads as zero here and their page count tells the truer story.

What remains

The unclaimed backlog, ranked by citation payoff. The ★ marking in the source sheet picks out exactly the Vedic saṃhitā / brāhmaṇa / upaniṣad / śrauta- and gṛhya-sūtra / prātiśākhya group.

How to read: One bar per unclaimed work, longest citation count first. Example 1: A short bar over a large page count is a low-yield, high-effort target. Example 2: The colour split shows how much of what is left is Vedic — the material with the awkward reference schemes.

Conclusion: The remaining work is Vedic and long. The kāvya and kośa material, which indexes quickly and is cited densely, is finished — so the backlog's citation payoff per page is the lowest the campaign has faced.

Source

← back to overview