CSL Observatory 13 years of Cologne Digital Sanskrit Lexicon

Created: 03-07-2026 · Last updated: 18-09-2026

Thirteen years, one scroll

The Cologne Digital Sanskrit Lexicon (CDSL) began in 1994, two decades before its public GitHub record. This page tells the narrower story that the committed 2014–2026 correction and repository snapshots can support: documented dictionary edits, issues, and commits. Earlier institutional and email history requires different evidence. Read it start to finish; it takes about seven minutes, in two chapters: the thirteen-year arc (told as a sticky-figure scrolly — one chart panel that changes as you scroll), then the first fully-measured month under the observatory.

Between and , the project logged individual, reconstructable corrections to dictionaries — a public ledger of philological repair with no real parallel in Sanskrit lexicography. This is what those thirteen years look like.

Corrections recorded

–

Dictionaries touched

from Apte to Böhtlingk-Roth

Hands on the work

correctors over 13 years

Thirteen years in one axis

Corrections recorded per year, → : documented edits across dictionaries, every one reconstructable from the committed snapshots. Every turning point in the steps below is a bend in this one series.

2014–2016 · The public correction ledger expands

The GitHub-era record opens with text repair, but this is not the project's founding: CDSL had already existed for twenty years. The inherited cfr.tsv form-correction file records fixes to transcription errors in the dictionary text: a ṭ read as द, a dropped conjunct, a mis-segmented compound. These are the form-layer corrections, and they dominate the early public ledger: the series peaks in , before the later git-derived layer begins.

What this proves: the earliest period represented by this dataset is dominated by form corrections. It does not prove what dominated the unmeasured 1994–2013 project history.

2019 · Pull requests arrive

The organisation and its public issue history date from 2014. 2019 is the narrower milestone when pull requests first appear in the committed snapshot. The number of distinct Git author identities active that year is . This marks adoption of an additional review mechanism, not the arrival of Git itself and not the beginning of public correction work.

The git era at full stretch

Through the campaign years the ledger surges as the git-era workflow lets several dictionaries be reworked in parallel. is the era's high-water mark: distinct authors active and commits — more than any year before it. The contributor base has since widened further to authors in , with commits as the observatory itself began committing data.

What this proves: the project's ceiling is a dozen people, not a hundred. Even at its most active it was a small circle working intensively — an important fact when reading everything that follows about concentration and continuity.

2025 · The correction wave, and the reckoning

2025 is the year the backlog crested. Issues were opened in bulk — of them, far more than any prior year — largely as a tracking mechanism for a fresh correction campaign, while closings lagged. The open-issue count carried into the next year reached its all-time high of . The project had, in effect, catalogued how much unfinished work it was actually carrying.

2026 · Taxonomy, and the observatory

The response was to start measuring. 2026 is the taxonomy-and-observatory era: a shared issue taxonomy pushed org-wide (pooled conformance now %), and closings finally outpacing openings — issues closed against opened — dropping the backlog from its peak to . This observatory is itself a product of that era: the project turning its own thirteen-year record into citable, reproducible data.

The harder facts

The arc above is the encouraging reading. But the same record carries four harder facts, and an honest story has to state them.

The work rests on one person. A single contributor — — accounts for % of every recorded contribution in the organisation. That is not a criticism of anyone; it is a structural risk.

Most of the backlog was never answered. Of the open issues still on the books, have never received a single reply — not a triage label, not a comment, nothing. Silence, not disagreement, is the dominant failure mode.

Issues that survive early tend to survive forever. The 2014 cohort makes the point starkly: % of the issues opened that year were still open four years later. Once an issue clears its first weeks unresolved, its odds of ever being closed collapse.

One thing did get fixed: licensing. When the observatory surfaced that 41 of the org's repositories carried no license at all, the project acted; after the RH1 rollout only remain unlicensed. A measured problem became a closed one.

Where a new contributor starts

If this story leaves you wanting to help rather than only to cite, the most valuable thing you can do is the least glamorous: answer a silent issue. The never-answered open issues are where a single reply — a triage label, a clarifying question, a "this is fixed" — has the highest marginal value in the entire organisation.

  • Triage the silence — the Issue Lifecycle and Taxonomy Triage pages surface the unanswered and unlabelled backlog, repo by repo.
  • See where the work is — the Ops Command view ranks repositories by open pressure and metadata blockers, so a first contribution lands where it counts.
  • Reuse the data — every figure on this page is downloadable from the Data page under CC-BY-4.0; the error-typology corpus is a published language resource in its own right.

Thirteen years of one small circle's careful work are now legible, citable, and open. The next chapter is whether that circle widens.

Trust block. Source: data/obs_t_timeline.csv, a committed feed loaded at build time. n = year × layer × component rows, summing to corrections. Data date: snapshot . Download: obs_t_timeline.csv.
Trust block. Source: data/obs_t_timeline.csv, a committed feed loaded at build time. n = year × layer × component rows; form era = years ≤ 2016. Data date: snapshot . Download: obs_t_timeline.csv.
Trust block. Source: data/velocity_timeline.csv, a committed feed loaded at build time. n = yearly rows, 2014–2026. Data date: snapshot . Download: velocity_timeline.csv.
Trust block. Source: data/velocity_timeline.csv, a committed feed loaded at build time. n = yearly rows (commits + distinct active authors per year). Data date: snapshot . Download: velocity_timeline.csv.
Trust block. Source: data/velocity_timeline.csv, a committed feed loaded at build time. n = yearly rows, 2014–2026. Data date: snapshot . Download: velocity_timeline.csv.
Trust block. Source: data/taxonomy_adoption.csv, a committed feed loaded at build time. n = yearly rows (pct_conformant per year; dashed line = all years pooled). Data date: snapshot . Download: taxonomy_adoption.csv.
Bus factor. contributor_identity.csv · n = contributors.
Silence. red = never answered ( of ). issue_lifecycle_backlog.csv · n = buckets.
Survival, 2014 cohort. issue_lifecycle_survival.csv · n = horizon rows.
Licensing. today's count from repo_health.csv · n = repos; 41 = the pre-rollout RH1 finding.

Answer a silent issue. The never-answered open issues are where one reply has the highest marginal value in the entire organisation.

  • Issue Lifecycle — the unanswered backlog, repo by repo
  • Taxonomy Triage — the unlabelled backlog
  • Ops Command — where the work is, ranked by open pressure
  • Data — every figure on this page, downloadable under CC-BY-4.0

Chapter 2 · The first measured month — July 2026

Chapter 1 ended with the observatory built and a loop promised: surface a fact, act on it, re-measure. This chapter is that loop's first full month of output. In July 2026 two scheduled data refreshes landed (July 20, July 28), and every result below was committed, with its dataset, inside the month it describes.

Commits in July

month # of in the whole record

Issues closed vs opened

/ the backlog kept draining

Repositories active

of ~85 in the organisation

The month in the thirteen-year curve

Chapter 1's closing claim — that 2026 is the year the ledger came under active management — holds at month granularity. July's commits make it the second-busiest month in the entire 2014–2026 record (only May 2026's was higher), and closings again outran openings, to .

How to read: Each bar is one month of org-wide commits, the twelve months to July 2026. Example 1: the three-bar plateau on the right — May, June, July 2026 each above 1,350 — is the observatory era running at a sustained pace no earlier period reached. Example 2: the low bars of late 2025 are the pre-taxonomy baseline the era is measured against.

:::note Trust block. Source: data/timeseries_monthly.csv, a committed feed loaded at build time. n = repo × month rows ( months shown). Data date: snapshot . Download: timeseries_monthly.csv. :::

Data table (figure fallback).

What this proves: the 2026 turn described in chapter 1 is not a single burst — it is a sustained operating pace, and July is the first month whose whole shape the observatory captured as it happened.

A volunteer campaign becomes a dataset

The single largest July commit was historiographical: the 2025–26 PWG scan-index campaign — eighteen months of volunteers page-indexing the printed editions that the Böhtlingk-Roth dictionary cites — was committed as data (PR #107): a registry of tracked works, of which are done, covering % of the tracked citation mass — pages indexed by volunteers. The full analysis is in the campaign report and on its dedicated page.

How to read: Each bar is one campaign status; length is the citation mass of the works in it, so the chart shows payoff at stake, not headcount. page-wise and nr-* are rulings, not backlog — works the campaign deliberately declined to index per-entry.

:::note Trust block. Source: data/pwg_scan_index.csv + data/pwg_scan_index_summary.json, committed feeds loaded at build time. n = tracked works. Data date: sheet snapshot . Report: reports/pwg_scan_index.md. Download: pwg_scan_index.csv. :::

Data table (figure fallback).

The campaign's follow-on question — can the indexed scans become e-text? — got a measured answer the same month, and the answer was no, not the obvious way. The kośa e-text pilot (PR #110) OCR'd the two heaviest kośa scan sets locally and recovered only 17.8% valid Sanskrit tokens — while the hOCR the Bayerische Staatsbibliothek already publishes for the same pages scored 43.8%, 2.5× better, for free. The job was re-scoped from OCR-from-scratch to ingest-and-correct, and the first BSB hOCR harvest landed on July 28 (PR #123). A NO-GO that costs one pilot and saves a campaign is the observatory loop working as designed.

What this proves: volunteer work that lived in a Google Sheet is now a versioned, downloadable dataset with its own regression checks — and the next step after it was chosen by measurement, not enthusiasm.

How many errors are left — the first population estimates

Thirteen years of corrections beg the question chapter 1 could not answer: how much is left? July produced the org's first defensible estimates (PR #120), using two-era Chapman capture–recapture over a measured record-linkage ladder: the form-era (2014–2019) and git-era (2019–2026) correction sets act as two capture occasions, and their overlap sizes the unseen population.

How to read: One row per dictionary with an estimable overlap; the dot is the Chapman point estimate of total error sites, the line its 95% confidence interval. Example 1: pw's estimate near 68,000 against ~11,000 sites already corrected implies most of the error population is still untouched. Example 2: a wide interval (bur) is honesty, not weakness — the overlap is small, and the method says so.

:::note Trust block. Source: data/error_recapture.csv, a committed feed loaded at build time. n = dictionaries with an estimable two-era overlap (of tested). Data date: 2026-07 (H1477 measurement; snapshot ). Estimator: two-era Chapman, capped at each dictionary's physical record count; assumptions and caveats in reports/error_recapture.md. Download: error_recapture.csv. :::

Data table (figure fallback).

The headline: pw has an estimated ~67,866 error sites (CI 59,208–76,525), of which ~56,935 are still uncorrected; mw ~60,997, with ~54,110 remaining. A within-era cross-check (PR #122) — correctors as capture occasions instead of eras — lands in the same order of magnitude and gives pwg its first estimate (~26,515). Just as important is what the linkage work rejected: the documented dead ends include the tempting <L>-number join, unsafe because 64% of form-era L-codes have drifted.

What this proves: the correction project's remaining work is now a number with a confidence interval, not a shrug. At the observed pace, the backlog of dictionary errors is measured in decades — which is exactly the kind of fact a funder or successor institution needs stated plainly.

Who actually cites the digital resource

Chapter 1 counted the work; July counted the audience. A systematic OpenAlex citation sweep (PR #128) replaced the five hand-picked citations the project used to show with a documented lower bound: scholarly works demonstrably naming or citing the Cologne digital lexicon — confirmed by name-unique phrase match, probable via the citation graph. The -work print envelope (works citing Monier-Williams or Böhtlingk in print) is deliberately fenced off, not claimed as digital reach, and the rejected phrase collisions stay in the CSV so the exclusions are evidenced, not asserted.

How to read: One bar per confidence tier of the sweep, external works only (project self-records excluded). Only the two green-to-amber bars are claimed as citations of the digital resource; the two grey bars are the honesty apparatus — the envelope not claimed, and the collisions rejected.

:::note Trust block. Source: data/citation_sweep.csv, a committed feed loaded at build time. n = candidate works from the OpenAlex sweep ( external after removing project self-records). Data date: API fetch 2026-07-28, cache committed. Method + recall bounds: reports/citation_sweep.md. Download: citation_sweep.csv. :::

Data table (figure fallback).

What this proves: the project can now put a sourced, reproducible number on its scholarly reach — smaller than a hand-wave, but real, and with its recall limits stated in the report rather than hidden.

The paper track, quietly

The same month moved the research pipeline without a single new figure needing to be drawn here. Blind cross-model double annotation put the error-typology corpus's inter-annotator agreement at κ = 0.906 [0.872–0.938] on the location axis (PR #102), clearing the gate the OBS-T paper had been waiting on. Its two rival manuscript drafts were reconciled into one canonical text (PR #125). A false Zenodo DOI was hunted down and corrected everywhere it had been asserted (PR #99); correction events got a persistent ID scheme so future releases stay comparable (PR #109); and the rights question was closed in the open — everything here publishes (PR #111).

What this proves: the loop chapter 1 promised — surface, act, re-measure — ran at monthly cadence for the first time in July 2026. Chapter 3 is whichever month next earns one.


Every figure on this page is computed from the committed datasets — snapshot . The scheduled refresh is intended to be monthly but may lag; the snapshot date, not the page-view date, is authoritative. Download the underlying CSV/JSON from the Data page to check any number here.

Dr. Mārcis Gasūns