Created: 03-07-2026 · Last updated: 18-09-2026
Thirteen years, one scroll
The Cologne Digital Sanskrit Lexicon (CDSL) began in 1994, two decades before its public GitHub record. This page tells the narrower story that the committed 2014–2026 correction and repository snapshots can support: documented dictionary edits, issues, and commits. Earlier institutional and email history requires different evidence. Read it start to finish; it takes about seven minutes, in two chapters: the thirteen-year arc (told as a sticky-figure scrolly — one chart panel that changes as you scroll), then the first fully-measured month under the observatory.
Between
Corrections recorded
Dictionaries touched
Hands on the work
Thirteen years in one axis
Corrections recorded per year,
2014–2016 · The public correction ledger expands
The GitHub-era record opens with text repair, but this is not the project's founding: CDSL had already existed for twenty years. The inherited cfr.tsv form-correction file records fixes to transcription errors in the dictionary text: a ṭ read as द, a dropped conjunct, a mis-segmented compound. These are the form-layer corrections, and they dominate the early public ledger: the series peaks in
What this proves: the earliest period represented by this dataset is dominated by form corrections. It does not prove what dominated the unmeasured 1994–2013 project history.
2019 · Pull requests arrive
The organisation and its public issue history date from 2014. 2019 is the narrower milestone when pull requests first appear in the committed snapshot. The number of distinct Git author identities active that year is
The git era at full stretch
Through the campaign years the ledger surges as the git-era workflow lets several dictionaries be reworked in parallel.
What this proves: the project's ceiling is a dozen people, not a hundred. Even at its most active it was a small circle working intensively — an important fact when reading everything that follows about concentration and continuity.
2025 · The correction wave, and the reckoning
2025 is the year the backlog crested. Issues were opened in bulk —
2026 · Taxonomy, and the observatory
The response was to start measuring. 2026 is the taxonomy-and-observatory era: a shared issue taxonomy pushed org-wide (pooled conformance now
The harder facts
The arc above is the encouraging reading. But the same record carries four harder facts, and an honest story has to state them.
The work rests on one person. A single contributor —
Most of the backlog was never answered. Of the open issues still on the books,
Issues that survive early tend to survive forever. The 2014 cohort makes the point starkly:
One thing did get fixed: licensing. When the observatory surfaced that 41 of the org's repositories carried no license at all, the project acted; after the RH1 rollout only
Where a new contributor starts
If this story leaves you wanting to help rather than only to cite, the most valuable thing you can do is the least glamorous: answer a silent issue. The
- Triage the silence — the Issue Lifecycle and Taxonomy Triage pages surface the unanswered and unlabelled backlog, repo by repo.
- See where the work is — the Ops Command view ranks repositories by open pressure and metadata blockers, so a first contribution lands where it counts.
- Reuse the data — every figure on this page is downloadable from the Data page under CC-BY-4.0; the error-typology corpus is a published language resource in its own right.
Thirteen years of one small circle's careful work are now legible, citable, and open. The next chapter is whether that circle widens.
data/obs_t_timeline.csv, a committed feed loaded at build time. n = obs_t_timeline.csv.data/obs_t_timeline.csv, a committed feed loaded at build time. n = obs_t_timeline.csv.data/velocity_timeline.csv, a committed feed loaded at build time. n = velocity_timeline.csv.data/velocity_timeline.csv, a committed feed loaded at build time. n = velocity_timeline.csv.data/velocity_timeline.csv, a committed feed loaded at build time. n = velocity_timeline.csv.data/taxonomy_adoption.csv, a committed feed loaded at build time. n = taxonomy_adoption.csv.contributor_identity.csv · n = issue_lifecycle_backlog.csv · n = issue_lifecycle_survival.csv · n = repo_health.csv · n = Answer a silent issue. The
- Issue Lifecycle — the unanswered backlog, repo by repo
- Taxonomy Triage — the unlabelled backlog
- Ops Command — where the work is, ranked by open pressure
- Data — every figure on this page, downloadable under CC-BY-4.0
Chapter 2 · The first measured month — July 2026
Chapter 1 ended with the observatory built and a loop promised: surface a fact, act on it, re-measure. This chapter is that loop's first full month of output. In July 2026 two scheduled data refreshes landed (July 20, July 28), and every result below was committed, with its dataset, inside the month it describes.
Commits in July
Issues closed vs opened
Repositories active
The month in the thirteen-year curve
Chapter 1's closing claim — that 2026 is the year the ledger came under active management — holds at month granularity. July's
How to read: Each bar is one month of org-wide commits, the twelve months to July 2026. Example 1: the three-bar plateau on the right — May, June, July 2026 each above 1,350 — is the observatory era running at a sustained pace no earlier period reached. Example 2: the low bars of late 2025 are the pre-taxonomy baseline the era is measured against.
:::note
Trust block. Source: data/timeseries_monthly.csv, a committed feed loaded at build time. n = timeseries_monthly.csv.
:::
Data table (figure fallback).
What this proves: the 2026 turn described in chapter 1 is not a single burst — it is a sustained operating pace, and July is the first month whose whole shape the observatory captured as it happened.
A volunteer campaign becomes a dataset
The single largest July commit was historiographical: the 2025–26 PWG scan-index campaign — eighteen months of volunteers page-indexing the printed editions that the Böhtlingk-Roth dictionary cites — was committed as data (PR #107): a registry of
How to read: Each bar is one campaign status; length is the citation mass of the works in it, so the chart shows payoff at stake, not headcount.
page-wiseandnr-*are rulings, not backlog — works the campaign deliberately declined to index per-entry.
:::note
Trust block. Source: data/pwg_scan_index.csv + data/pwg_scan_index_summary.json, committed feeds loaded at build time. n = reports/pwg_scan_index.md. Download: pwg_scan_index.csv.
:::
Data table (figure fallback).
The campaign's follow-on question — can the indexed scans become e-text? — got a measured answer the same month, and the answer was no, not the obvious way. The kośa e-text pilot (PR #110) OCR'd the two heaviest kośa scan sets locally and recovered only 17.8% valid Sanskrit tokens — while the hOCR the Bayerische Staatsbibliothek already publishes for the same pages scored 43.8%, 2.5× better, for free. The job was re-scoped from OCR-from-scratch to ingest-and-correct, and the first BSB hOCR harvest landed on July 28 (PR #123). A NO-GO that costs one pilot and saves a campaign is the observatory loop working as designed.
What this proves: volunteer work that lived in a Google Sheet is now a versioned, downloadable dataset with its own regression checks — and the next step after it was chosen by measurement, not enthusiasm.
How many errors are left — the first population estimates
Thirteen years of corrections beg the question chapter 1 could not answer: how much is left? July produced the org's first defensible estimates (PR #120), using two-era Chapman capture–recapture over a measured record-linkage ladder: the form-era (2014–2019) and git-era (2019–2026) correction sets act as two capture occasions, and their overlap sizes the unseen population.
How to read: One row per dictionary with an estimable overlap; the dot is the Chapman point estimate of total error sites, the line its 95% confidence interval. Example 1: pw's estimate near 68,000 against ~11,000 sites already corrected implies most of the error population is still untouched. Example 2: a wide interval (bur) is honesty, not weakness — the overlap is small, and the method says so.
:::note
Trust block. Source: data/error_recapture.csv, a committed feed loaded at build time. n = reports/error_recapture.md. Download: error_recapture.csv.
:::
Data table (figure fallback).
The headline: pw has an estimated ~67,866 error sites (CI 59,208–76,525), of which ~56,935 are still uncorrected; mw ~60,997, with ~54,110 remaining. A within-era cross-check (PR #122) — correctors as capture occasions instead of eras — lands in the same order of magnitude and gives pwg its first estimate (~26,515). Just as important is what the linkage work rejected: the documented dead ends include the tempting <L>-number join, unsafe because 64% of form-era L-codes have drifted.
What this proves: the correction project's remaining work is now a number with a confidence interval, not a shrug. At the observed pace, the backlog of dictionary errors is measured in decades — which is exactly the kind of fact a funder or successor institution needs stated plainly.
Who actually cites the digital resource
Chapter 1 counted the work; July counted the audience. A systematic OpenAlex citation sweep (PR #128) replaced the five hand-picked citations the project used to show with a documented lower bound:
How to read: One bar per confidence tier of the sweep, external works only (project self-records excluded). Only the two green-to-amber bars are claimed as citations of the digital resource; the two grey bars are the honesty apparatus — the envelope not claimed, and the collisions rejected.
:::note
Trust block. Source: data/citation_sweep.csv, a committed feed loaded at build time. n = reports/citation_sweep.md. Download: citation_sweep.csv.
:::
Data table (figure fallback).
What this proves: the project can now put a sourced, reproducible number on its scholarly reach — smaller than a hand-wave, but real, and with its recall limits stated in the report rather than hidden.
The paper track, quietly
The same month moved the research pipeline without a single new figure needing to be drawn here. Blind cross-model double annotation put the error-typology corpus's inter-annotator agreement at κ = 0.906 [0.872–0.938] on the location axis (PR #102), clearing the gate the OBS-T paper had been waiting on. Its two rival manuscript drafts were reconciled into one canonical text (PR #125). A false Zenodo DOI was hunted down and corrected everywhere it had been asserted (PR #99); correction events got a persistent ID scheme so future releases stay comparable (PR #109); and the rights question was closed in the open — everything here publishes (PR #111).
What this proves: the loop chapter 1 promised — surface, act, re-measure — ran at monthly cadence for the first time in July 2026. Chapter 3 is whichever month next earns one.
Every figure on this page is computed from the committed datasets — snapshot
Dr. Mārcis Gasūns