CSL Observatory 13 years of Cologne Digital Sanskrit Lexicon

Thirteen years, one scroll

The Cologne Digital Sanskrit Lexicon (CDSL) began in 1994, two decades before its public GitHub record. This page tells the narrower story that the committed 2014–2026 correction and repository snapshots can support: documented dictionary edits, issues, and commits. Earlier institutional and email history requires different evidence. Read it start to finish; it takes about seven minutes, in two chapters: the thirteen-year arc, then the first fully-measured month under the observatory.

Between and , the project logged individual, reconstructable corrections to dictionaries — a public ledger of philological repair with no real parallel in Sanskrit lexicography. This is what those thirteen years look like.

Corrections recorded

Dictionaries touched

from Apte to Böhtlingk-Roth

Hands on the work

correctors over 13 years

The spine: a backlog that tells the whole story

The single most honest summary of the project's history is the shape of its open-issue backlog — the count of unresolved issues carried into each year. It rises through the campaign years, holds, and then, in 2026, drops sharply as the taxonomy-and-observatory era brings the ledger under active management. Every turning point in the prose below is a bend in this one line.

:::note Trust block. Source: data/velocity_timeline.csv, a committed feed loaded at build time. n = yearly rows, 2014–2026. Data date: snapshot . Download: velocity_timeline.csv. :::

Data table (figure fallback).

What this proves: the project has always generated far more work than any small team could close, and the backlog is the accumulated evidence. What changed in 2026 is not that the work got smaller — it is that the org finally began measuring and draining it.

2014–2016 · The public correction ledger expands

The GitHub-era record opens with text repair, but this is not the project's founding: CDSL had already existed for twenty years. The inherited cfr.tsv form-correction file records fixes to transcription errors in the dictionary text: a read as , a dropped conjunct, a mis-segmented compound. These are the form-layer corrections, and they dominate the early public ledger: the series peaks in 2015–2016, before the later git-derived layer begins.

:::note Trust block. Source: data/obs_t_timeline.csv, a committed feed loaded at build time. n = year × layer × component rows, summing to corrections. Data date: snapshot . Download: obs_t_timeline.csv. :::

Data table (figure fallback).

What this proves: the earliest period represented by this dataset is dominated by form corrections. It does not prove what dominated the unmeasured 1994–2013 project history.

2019 · Pull requests arrive

The organisation and its public issue history date from 2014. 2019 is the narrower milestone when pull requests first appear in the committed snapshot. The number of distinct Git author identities active that year is . This marks adoption of an additional review mechanism, not the arrival of Git itself and not the beginning of public correction work.

2021 · The volume peak

If any single year was the project at full stretch, it was : distinct authors active — the widest the contributor base has ever been — and commits, more than any year before it. The correction ledger surges again as the git-era workflow lets several dictionaries be reworked in parallel.

Peak breadth ()

distinct active authors

Commits that year

an all-time high

Issues opened

campaign in full flow

What this proves: the project's ceiling is a dozen people, not a hundred. Even at its most active it was a small circle working intensively — an important fact when reading everything that follows about concentration and continuity.

2025 · The correction wave, and the reckoning

2025 is the year the backlog crested. Issues were opened in bulk — of them, far more than any prior year — largely as a tracking mechanism for a fresh correction campaign, while closings lagged. The open-issue count carried into the next year reached its all-time high of . The project had, in effect, catalogued how much unfinished work it was actually carrying.

2026 · Taxonomy, and the observatory

The response was to start measuring. 2026 is the taxonomy-and-observatory era: a shared issue taxonomy pushed org-wide (pooled conformance now %), and closings finally outpacing openings — issues closed against opened — dropping the backlog from its 1,742 peak to . This observatory is itself a product of that era: the project turning its own thirteen-year record into citable, reproducible data.


The arc above is the encouraging reading. But the same record carries four harder facts, and an honest story has to state them.

The work rests on one person

Across all thirteen years, a single contributor — — accounts for % of every recorded contribution in the organisation. That is not a criticism of anyone; it is a structural risk. A project this concentrated is one departure away from stalling, and no amount of tooling changes that. It is the first thing a would-be funder, host institution, or successor needs to know. Community analysis →

Most of the backlog was never answered

Of the open issues still on the books, have never received a single reply — not a triage label, not a comment, nothing. Silence, not disagreement, is the dominant failure mode: work is filed and then quietly outlives everyone's attention. The backlog is not a queue being worked down in order; it is a sediment, and most of it has never been touched since the day it was opened. Issue lifecycle →

Issues that survive early tend to survive forever

The 2014 cohort makes the point starkly: % of the issues opened that year were still open four years later. Once an issue clears its first weeks unresolved, its odds of ever being closed collapse. This is why the -issue silence matters — the backlog does not decay on its own; unattended issues become permanent.

:::note Trust block. Source: data/issue_lifecycle_survival.csv, a committed feed loaded at build time. n = horizon rows for the 2014 cohort (of cohort × horizon rows). Data date: snapshot . Download: issue_lifecycle_survival.csv. :::

Data table (figure fallback).

What this proves: the 2014 cohort's survival curve flattens well above zero — it never approaches full resolution. An issue's fate is largely sealed in its first months.

One thing did get fixed: licensing

The record is not only decline. When the observatory surfaced that 41 of the org's repositories carried no license at all — a FAIR-reuse violation that made the data legally unsafe to build on — the project acted. After the RH1 license rollout, only repositories remain unlicensed, and those are the archive candidates intentionally held back for a separate cleanup. A measured problem became a closed one. Repository health →

What this proves: the observatory is not a mirror the project looks into and sighs at — the licensing repair (41 → ) is the template. Surface a fact, act on it, re-measure. That is the loop this whole site exists to enable.

Where a new contributor starts

If this story leaves you wanting to help rather than only to cite, the most valuable thing you can do is the least glamorous: answer a silent issue. The never-answered open issues are where a single reply — a triage label, a clarifying question, a "this is fixed" — has the highest marginal value in the entire organisation.

Thirteen years of one small circle's careful work are now legible, citable, and open. The next chapter is whether that circle widens.


Chapter 2 · The first measured month — July 2026

Chapter 1 ended with the observatory built and a loop promised: surface a fact, act on it, re-measure. This chapter is that loop's first full month of output. In July 2026 two scheduled data refreshes landed (July 20, July 28), and every result below was committed, with its dataset, inside the month it describes.

Commits in July

month # of in the whole record

Issues closed vs opened

/ the backlog kept draining

Repositories active

of ~85 in the organisation

The month in the thirteen-year curve

Chapter 1's closing claim — that 2026 is the year the ledger came under active management — holds at month granularity. July's commits make it the second-busiest month in the entire 2014–2026 record (only May 2026's was higher), and closings again outran openings, to .

How to read: Each bar is one month of org-wide commits, the twelve months to July 2026. Example 1: the three-bar plateau on the right — May, June, July 2026 each above 1,350 — is the observatory era running at a sustained pace no earlier period reached. Example 2: the low bars of late 2025 are the pre-taxonomy baseline the era is measured against.

:::note Trust block. Source: data/timeseries_monthly.csv, a committed feed loaded at build time. n = repo × month rows ( months shown). Data date: snapshot . Download: timeseries_monthly.csv. :::

Data table (figure fallback).

What this proves: the 2026 turn described in chapter 1 is not a single burst — it is a sustained operating pace, and July is the first month whose whole shape the observatory captured as it happened.

A volunteer campaign becomes a dataset

The single largest July commit was historiographical: the 2025–26 PWG scan-index campaign — eighteen months of volunteers page-indexing the printed editions that the Böhtlingk-Roth dictionary cites — was committed as data (PR #107): a registry of tracked works, of which are done, covering % of the tracked citation mass — pages indexed by volunteers. The full analysis is in the campaign report and on its dedicated page.

How to read: Each bar is one campaign status; length is the citation mass of the works in it, so the chart shows payoff at stake, not headcount. page-wise and nr-* are rulings, not backlog — works the campaign deliberately declined to index per-entry.

:::note Trust block. Source: data/pwg_scan_index.csv + data/pwg_scan_index_summary.json, committed feeds loaded at build time. n = tracked works. Data date: sheet snapshot . Report: reports/pwg_scan_index.md. Download: pwg_scan_index.csv. :::

Data table (figure fallback).

The campaign's follow-on question — can the indexed scans become e-text? — got a measured answer the same month, and the answer was no, not the obvious way. The kośa e-text pilot (PR #110) OCR'd the two heaviest kośa scan sets locally and recovered only 17.8% valid Sanskrit tokens — while the hOCR the Bayerische Staatsbibliothek already publishes for the same pages scored 43.8%, 2.5× better, for free. The job was re-scoped from OCR-from-scratch to ingest-and-correct, and the first BSB hOCR harvest landed on July 28 (PR #123). A NO-GO that costs one pilot and saves a campaign is the observatory loop working as designed.

What this proves: volunteer work that lived in a Google Sheet is now a versioned, downloadable dataset with its own regression checks — and the next step after it was chosen by measurement, not enthusiasm.

How many errors are left — the first population estimates

Thirteen years of corrections beg the question chapter 1 could not answer: how much is left? July produced the org's first defensible estimates (PR #120), using two-era Chapman capture–recapture over a measured record-linkage ladder: the form-era (2014–2019) and git-era (2019–2026) correction sets act as two capture occasions, and their overlap sizes the unseen population.

How to read: One row per dictionary with an estimable overlap; the dot is the Chapman point estimate of total error sites, the line its 95% confidence interval. Example 1: pw's estimate near 68,000 against ~11,000 sites already corrected implies most of the error population is still untouched. Example 2: a wide interval (bur) is honesty, not weakness — the overlap is small, and the method says so.

:::note Trust block. Source: data/error_recapture.csv, a committed feed loaded at build time. n = dictionaries with an estimable two-era overlap (of tested). Data date: 2026-07 (H1477 measurement; snapshot ). Estimator: two-era Chapman, capped at each dictionary's physical record count; assumptions and caveats in reports/error_recapture.md. Download: error_recapture.csv. :::

Data table (figure fallback).

The headline: pw has an estimated ~67,866 error sites (CI 59,208–76,525), of which ~56,935 are still uncorrected; mw ~60,997, with ~54,110 remaining. A within-era cross-check (PR #122) — correctors as capture occasions instead of eras — lands in the same order of magnitude and gives pwg its first estimate (~26,515). Just as important is what the linkage work rejected: the documented dead ends include the tempting <L>-number join, unsafe because 64% of form-era L-codes have drifted.

What this proves: the correction project's remaining work is now a number with a confidence interval, not a shrug. At the observed pace, the backlog of dictionary errors is measured in decades — which is exactly the kind of fact a funder or successor institution needs stated plainly.

Who actually cites the digital resource

Chapter 1 counted the work; July counted the audience. A systematic OpenAlex citation sweep (PR #128) replaced the five hand-picked citations the project used to show with a documented lower bound: scholarly works demonstrably naming or citing the Cologne digital lexicon — confirmed by name-unique phrase match, probable via the citation graph. The -work print envelope (works citing Monier-Williams or Böhtlingk in print) is deliberately fenced off, not claimed as digital reach, and the rejected phrase collisions stay in the CSV so the exclusions are evidenced, not asserted.

How to read: One bar per confidence tier of the sweep, external works only (project self-records excluded). Only the two green-to-amber bars are claimed as citations of the digital resource; the two grey bars are the honesty apparatus — the envelope not claimed, and the collisions rejected.

:::note Trust block. Source: data/citation_sweep.csv, a committed feed loaded at build time. n = candidate works from the OpenAlex sweep ( external after removing project self-records). Data date: API fetch 2026-07-28, cache committed. Method + recall bounds: reports/citation_sweep.md. Download: citation_sweep.csv. :::

Data table (figure fallback).

What this proves: the project can now put a sourced, reproducible number on its scholarly reach — smaller than a hand-wave, but real, and with its recall limits stated in the report rather than hidden.

The paper track, quietly

The same month moved the research pipeline without a single new figure needing to be drawn here. Blind cross-model double annotation put the error-typology corpus's inter-annotator agreement at κ = 0.906 [0.872–0.938] on the location axis (PR #102), clearing the gate the OBS-T paper had been waiting on. Its two rival manuscript drafts were reconciled into one canonical text (PR #125). A false Zenodo DOI was hunted down and corrected everywhere it had been asserted (PR #99); correction events got a persistent ID scheme so future releases stay comparable (PR #109); and the rights question was closed in the open — everything here publishes (PR #111).

What this proves: the loop chapter 1 promised — surface, act, re-measure — ran at monthly cadence for the first time in July 2026. Chapter 3 is whichever month next earns one.


Every figure on this page is computed from the committed datasets — snapshot . The scheduled refresh is intended to be monthly but may lag; the snapshot date, not the page-view date, is authoritative. Download the underlying CSV/JSON from the Data page to check any number here.