CSL Observatory 13 years of Cologne Digital Sanskrit Lexicon

Correction Anatomy

Who fixes what in the OBS-T correction corpus, and how letters break: the character-confusion structure, its directionality, and the concentration and tenure of correction labor. All data is read from the committed OBS-T aggregates; the inferential companions are hypotheses H5, H6 and H8 in the Phase-3 spec.

Character Confusion Matrix

Every substitution edit in the corpus is a from → to character pair. This matrix shows the full character-level error space for one encoding regime at a time (IAST Sanskrit spans by default; raw markup/SLP1 edits are separate regimes and are not mixed in). The top 40 characters by total involvement get their own row/column; everything rarer is binned into "·other".

How to read: Row = the character that was wrong (from); column = what it was corrected to (to); colour = event count on a log scale, so both the dominant pairs and the long tail stay visible. Example 1: A bright cell at row P, column p means upper-case→lower-case folding of that letter is a dominant correction. Example 2: A bright diacritic cell such as m → ṃ marks anusvāra repair — a diacritic restoration, not a lexical change.

Conclusion: In the IAST space, case folding (P → p, 2,945 events) and punctuation swaps dominate the character-level error surface — corrections are typographic, not lexical. This is the OBS-T micro-edit finding (H1) made visible: the corpus's modal correction changes one character's case or diacritic, not a word.

Confusion Asymmetry

Companion to hypothesis H8: are confusions directional? For every unordered character pair with at least 30 total events in the selected edit space, this plots the share of events flowing in the pair's dominant direction against the pair's total volume. Under the null (symmetric noise) shares sit near 0.5; a systematic error source (OCR, case-folding, encoding conversion) pushes shares toward 1. Whiskers are Wilson 95% CIs; emphasized dots are significant after a Benjamini-Hochberg correction (exact two-sided binomial, q < 0.05) across all tested pairs.

How to read: Each dot is one character pair; x = total events (log), y = share in the dominant direction (0.5 = perfectly symmetric, 1.0 = one-way). Example 1: A large-volume pair at share ≈ 1.0 with a tight CI is a one-way error channel — e.g. upper case always folding down, never up. Example 2: A pair whose CI whisker crosses 0.5 (muted colour) is directionally undecided — its two directions occur at statistically indistinguishable rates.

Conclusion: The high-volume confusions sit far above 0.5 with tight CIs: they have a dominant direction, so the error source is systematic — OCR misreads, case-folding, and encoding conversion each push characters one way — rather than symmetric typing noise. That directionality is exploitable: a correction assistant can rank candidate fixes by the known flow direction (H8's practical payoff).

Corrector Pareto Curve

The correction corpus records 52k events by 60 attributed correctors. Ranking correctors by event count and accumulating their share shows how concentrated the labor is — the event-level restatement of the org's bus-factor finding.

How to read: x = correctors ranked by events (1 = most prolific), y = cumulative share of all events; the dashed line marks 80%. Example 1: The curve crossing 0.8 at rank 2 means two people carry four fifths of thirteen years of corrections. Example 2: A long flat tail past rank 10 means the remaining correctors are drive-by contributors whose combined weight is marginal.

The same head of the distribution, in absolute events:

Conclusion: Two of 60 correctors — funderburkjim (35,057 events) and drdhaval2785 (8,248) — carry ~83% of the corpus. The bus-factor risk measured elsewhere at repo level holds at the level of individual correction events: the corpus is, empirically, the work of two hands plus a long tail.

Corrector × Component Matrix

Companion to hypotheses H5/H6: do specialists exist within the corpus? Each row is one of the top-8 correctors; each cell is the share of that corrector's events landing in one error component (row-normalized, so prolific and occasional correctors are comparable).

How to read: Row = corrector, column = error component (the part of the dictionary entry repaired), colour = the share of that corrector's own events in that component. Example 1: A row with one very bright cell is a specialist — nearly all their corrections repair one component. Example 2: Rows with similar colour profiles mean the typology reflects the material, not the person — everyone fixes roughly the same mix (H5's invariance claim).

Conclusion: The matrix answers whether the corpus has internal division of labor. Broadly similar row profiles (sense-heavy for both core correctors) support H5's corrector-invariance claim — the error typology is a property of the dictionaries, not of who happened to fix them; sharply divergent rows would instead reveal in-corpus specialisation. Note the unattributed column reflects join failures, not a correction type.

Corrector Tenure Spans

Each corrector's active span, from their first to their last recorded correction event. Line thickness scales with the log of their event count.

How to read: Each horizontal line is one corrector (top 20 by events); it starts at their first event and ends at their last; thicker = more events. Example 1: A thick line spanning 2014–2026 is a founder-maintainer whose correction work never stopped. Example 2: A thin line a few weeks long is a drive-by corrector — one campaign, then gone.

Conclusion: Correction labor is long-tenure, few-hands: the core pair's spans run 2014–2026 while most of the top 20 are short, bounded engagements. Sustainability of the correction corpus therefore depends on the same two-person continuity the community pages flag at repo level.

Back to overview