Thirteen years, one scroll
The Cologne Digital Sanskrit Lexicon (CDSL) began in 1994, two decades before its public GitHub record. This page tells the narrower story that the committed 2014–2026 correction and repository snapshots can support: documented dictionary edits, issues, and commits. Earlier institutional and email history requires different evidence. Read it start to finish; it takes about seven minutes, in two chapters: the thirteen-year arc, then the first fully-measured month under the observatory.
Between
Corrections recorded
Dictionaries touched
Hands on the work
The spine: a backlog that tells the whole story
The single most honest summary of the project's history is the shape of its open-issue backlog — the count of unresolved issues carried into each year. It rises through the campaign years, holds, and then, in 2026, drops sharply as the taxonomy-and-observatory era brings the ledger under active management. Every turning point in the prose below is a bend in this one line.
:::note
Trust block. Source: data/velocity_timeline.csv, a committed feed loaded at build time. n = velocity_timeline.csv.
:::
Data table (figure fallback).
What this proves: the project has always generated far more work than any small team could close, and the backlog is the accumulated evidence. What changed in 2026 is not that the work got smaller — it is that the org finally began measuring and draining it.
2014–2016 · The public correction ledger expands
The GitHub-era record opens with text repair, but this is not the project's founding: CDSL had already existed for twenty years. The inherited cfr.tsv form-correction file records fixes to transcription errors in the dictionary text: a ṭ read as द, a dropped conjunct, a mis-segmented compound. These are the form-layer corrections, and they dominate the early public ledger: the series peaks in 2015–2016, before the later git-derived layer begins.
:::note
Trust block. Source: data/obs_t_timeline.csv, a committed feed loaded at build time. n = obs_t_timeline.csv.
:::
Data table (figure fallback).
What this proves: the earliest period represented by this dataset is dominated by form corrections. It does not prove what dominated the unmeasured 1994–2013 project history.
2019 · Pull requests arrive
The organisation and its public issue history date from 2014. 2019 is the narrower milestone when pull requests first appear in the committed snapshot. The number of distinct Git author identities active that year is
2021 · The volume peak
If any single year was the project at full stretch, it was
Peak breadth ( )
Commits that year
Issues opened
What this proves: the project's ceiling is a dozen people, not a hundred. Even at its most active it was a small circle working intensively — an important fact when reading everything that follows about concentration and continuity.
2025 · The correction wave, and the reckoning
2025 is the year the backlog crested. Issues were opened in bulk —
2026 · Taxonomy, and the observatory
The response was to start measuring. 2026 is the taxonomy-and-observatory era: a shared issue taxonomy pushed org-wide (pooled conformance now
The arc above is the encouraging reading. But the same record carries four harder facts, and an honest story has to state them.
The work rests on one person
Across all thirteen years, a single contributor —
Most of the backlog was never answered
Of the open issues still on the books,
Issues that survive early tend to survive forever
The 2014 cohort makes the point starkly:
:::note
Trust block. Source: data/issue_lifecycle_survival.csv, a committed feed loaded at build time. n = issue_lifecycle_survival.csv.
:::
Data table (figure fallback).
What this proves: the 2014 cohort's survival curve flattens well above zero — it never approaches full resolution. An issue's fate is largely sealed in its first months.
One thing did get fixed: licensing
The record is not only decline. When the observatory surfaced that 41 of the org's repositories carried no license at all — a FAIR-reuse violation that made the data legally unsafe to build on — the project acted. After the RH1 license rollout, only
What this proves: the observatory is not a mirror the project looks into and sighs at — the licensing repair (41 →
) is the template. Surface a fact, act on it, re-measure. That is the loop this whole site exists to enable.
Where a new contributor starts
If this story leaves you wanting to help rather than only to cite, the most valuable thing you can do is the least glamorous: answer a silent issue. The
- Triage the silence — the Issue Lifecycle and Taxonomy Triage pages surface the unanswered and unlabelled backlog, repo by repo.
- See where the work is — the Ops Command view ranks repositories by open pressure and metadata blockers, so a first contribution lands where it counts.
- Reuse the data — every figure on this page is downloadable from the Data page under CC-BY-4.0; the error-typology corpus is a published language resource in its own right.
Thirteen years of one small circle's careful work are now legible, citable, and open. The next chapter is whether that circle widens.
Chapter 2 · The first measured month — July 2026
Chapter 1 ended with the observatory built and a loop promised: surface a fact, act on it, re-measure. This chapter is that loop's first full month of output. In July 2026 two scheduled data refreshes landed (July 20, July 28), and every result below was committed, with its dataset, inside the month it describes.
Commits in July
Issues closed vs opened
Repositories active
The month in the thirteen-year curve
Chapter 1's closing claim — that 2026 is the year the ledger came under active management — holds at month granularity. July's
How to read: Each bar is one month of org-wide commits, the twelve months to July 2026. Example 1: the three-bar plateau on the right — May, June, July 2026 each above 1,350 — is the observatory era running at a sustained pace no earlier period reached. Example 2: the low bars of late 2025 are the pre-taxonomy baseline the era is measured against.
:::note
Trust block. Source: data/timeseries_monthly.csv, a committed feed loaded at build time. n = timeseries_monthly.csv.
:::
Data table (figure fallback).
What this proves: the 2026 turn described in chapter 1 is not a single burst — it is a sustained operating pace, and July is the first month whose whole shape the observatory captured as it happened.
A volunteer campaign becomes a dataset
The single largest July commit was historiographical: the 2025–26 PWG scan-index campaign — eighteen months of volunteers page-indexing the printed editions that the Böhtlingk-Roth dictionary cites — was committed as data (PR #107): a registry of
How to read: Each bar is one campaign status; length is the citation mass of the works in it, so the chart shows payoff at stake, not headcount.
page-wiseandnr-*are rulings, not backlog — works the campaign deliberately declined to index per-entry.
:::note
Trust block. Source: data/pwg_scan_index.csv + data/pwg_scan_index_summary.json, committed feeds loaded at build time. n = reports/pwg_scan_index.md. Download: pwg_scan_index.csv.
:::
Data table (figure fallback).
The campaign's follow-on question — can the indexed scans become e-text? — got a measured answer the same month, and the answer was no, not the obvious way. The kośa e-text pilot (PR #110) OCR'd the two heaviest kośa scan sets locally and recovered only 17.8% valid Sanskrit tokens — while the hOCR the Bayerische Staatsbibliothek already publishes for the same pages scored 43.8%, 2.5× better, for free. The job was re-scoped from OCR-from-scratch to ingest-and-correct, and the first BSB hOCR harvest landed on July 28 (PR #123). A NO-GO that costs one pilot and saves a campaign is the observatory loop working as designed.
What this proves: volunteer work that lived in a Google Sheet is now a versioned, downloadable dataset with its own regression checks — and the next step after it was chosen by measurement, not enthusiasm.
How many errors are left — the first population estimates
Thirteen years of corrections beg the question chapter 1 could not answer: how much is left? July produced the org's first defensible estimates (PR #120), using two-era Chapman capture–recapture over a measured record-linkage ladder: the form-era (2014–2019) and git-era (2019–2026) correction sets act as two capture occasions, and their overlap sizes the unseen population.
How to read: One row per dictionary with an estimable overlap; the dot is the Chapman point estimate of total error sites, the line its 95% confidence interval. Example 1: pw's estimate near 68,000 against ~11,000 sites already corrected implies most of the error population is still untouched. Example 2: a wide interval (bur) is honesty, not weakness — the overlap is small, and the method says so.
:::note
Trust block. Source: data/error_recapture.csv, a committed feed loaded at build time. n = reports/error_recapture.md. Download: error_recapture.csv.
:::
Data table (figure fallback).
The headline: pw has an estimated ~67,866 error sites (CI 59,208–76,525), of which ~56,935 are still uncorrected; mw ~60,997, with ~54,110 remaining. A within-era cross-check (PR #122) — correctors as capture occasions instead of eras — lands in the same order of magnitude and gives pwg its first estimate (~26,515). Just as important is what the linkage work rejected: the documented dead ends include the tempting <L>-number join, unsafe because 64% of form-era L-codes have drifted.
What this proves: the correction project's remaining work is now a number with a confidence interval, not a shrug. At the observed pace, the backlog of dictionary errors is measured in decades — which is exactly the kind of fact a funder or successor institution needs stated plainly.
Who actually cites the digital resource
Chapter 1 counted the work; July counted the audience. A systematic OpenAlex citation sweep (PR #128) replaced the five hand-picked citations the project used to show with a documented lower bound:
How to read: One bar per confidence tier of the sweep, external works only (project self-records excluded). Only the two green-to-amber bars are claimed as citations of the digital resource; the two grey bars are the honesty apparatus — the envelope not claimed, and the collisions rejected.
:::note
Trust block. Source: data/citation_sweep.csv, a committed feed loaded at build time. n = reports/citation_sweep.md. Download: citation_sweep.csv.
:::
Data table (figure fallback).
What this proves: the project can now put a sourced, reproducible number on its scholarly reach — smaller than a hand-wave, but real, and with its recall limits stated in the report rather than hidden.
The paper track, quietly
The same month moved the research pipeline without a single new figure needing to be drawn here. Blind cross-model double annotation put the error-typology corpus's inter-annotator agreement at κ = 0.906 [0.872–0.938] on the location axis (PR #102), clearing the gate the OBS-T paper had been waiting on. Its two rival manuscript drafts were reconciled into one canonical text (PR #125). A false Zenodo DOI was hunted down and corrected everywhere it had been asserted (PR #99); correction events got a persistent ID scheme so future releases stay comparable (PR #109); and the rights question was closed in the open — everything here publishes (PR #111).
What this proves: the loop chapter 1 promised — surface, act, re-measure — ran at monthly cadence for the first time in July 2026. Chapter 3 is whichever month next earns one.
Every figure on this page is computed from the committed datasets — snapshot