CSL Observatory 13 years of Cologne Digital Sanskrit Lexicon

Org Shape

The organization's labor map and weekly trajectory: who works where across the 78 sanskrit-lexicon repositories, how specialised each contributor is, and how the org-level backlog moves between weekly snapshots. Contributor tables were computed 2026-06 from the full commit history; snapshots accrue weekly via the refresh workflow. The inferential companions are hypotheses H4 and H7 in the Phase-3 spec.

Contributor × Repository Heatmap

The org's labor map: every contributor (rows) against every repository they have committed to (columns), grouped into repo families. Computed in June 2026 from the full commit history and never rendered until now.

How to read: Row = contributor login; column = repository, grouped by family panel; colour = commit count on a log scale. Example 1: A row lighting up across an entire family panel is a family owner — one person maintaining a whole class of repos. Example 2: A column with several coloured rows is a shared repo — the closest thing the org has to a commons; a column with exactly one row is a single-maintainer repo, the bus-factor surface.

Conclusion: The map shows where the org's 13 years of commits actually landed: a dense two-row core (funderburkjim, drdhaval2785) spanning nearly every family, a handful of family-scoped contributors, and a long tail of single-repo participants. Columns with a single coloured cell are the repos that stop moving if one person does.

Specialisation Index

Companion to hypothesis H4: is the org run by generalists or specialists? Normalized entropy of each contributor's commit distribution across repos — 0 means all commits in one repo (pure specialist), 1 means commits spread evenly over every repo they touch (pure generalist).

How to read: Each dot is one contributor; x = normalized entropy, dot size = total commits, colour = the repo family holding the largest share of their work. Example 1: A large dot at high entropy is a high-volume generalist — the org's infrastructure depends on their breadth. Example 2: A small dot near 0 is a focused contributor whose entire output lives in one repository.

Conclusion: The org runs on generalists: the highest-volume contributors sit at high entropy (Jim Funderburk at 0.69 across 58 repos, Dhaval Patel at 0.78 across 39), while true specialists are low-volume. That is H4's picture — labor is concentrated in people, not confined to repos — and it means the org's knowledge is portable across repos but not across people.

Dominant-Family Capture

How captured each contributor is by a single repo family: the share of their commits in their dominant family. The 0.5 rule marks the majority line — right of it, more than half of a contributor's work lives in one family.

How to read: Bar length = share of the contributor's commits in their dominant family; colour = which family that is; the dashed rule marks 50%. Example 1: A bar near 1.0 is a contributor wholly captured by one family — their departure affects exactly one class of repos. Example 2: A bar just past 0.5 with high total commits is a broad contributor with a lean — invested in one family but active elsewhere.

Conclusion: Most contributors sit right of the majority line — one family holds most of their work — but the core maintainers do not: their dominant-family share stays near 0.5–0.65 despite huge volume. Family capture is a property of the periphery; the center is diversified.

Snapshot Drift

Companion to hypothesis H7: is the org backlog in steady state? Weekly snapshots of org-level totals, shown as percent change from the first snapshot so all four metrics share one scale. H7 is registered but deliberately underpowered at the current snapshot count — with only five weekly snapshots, Mann-Kendall trend tests cannot reach significance except for perfect monotone runs — so this view is descriptive until ≥ 10 snapshots accumulate.

How to read: One panel per metric; x = snapshot date, y = percent change since the first snapshot. Example 1: A flat line at 0% is a metric in steady state — the H7 null holds descriptively. Example 2: A panel climbing week over week (e.g. total issues) means the backlog is growing faster than it is being drained.

The same series in absolute counts (log scale, one line per metric):

Conclusion: Until now the weekly snapshots existed only as a text digest; this is their first rendering. Movements so far are small — consistent with the steady-state null of H7 — but the verdict is explicitly deferred until enough snapshots accumulate for the registered Mann-Kendall test to have power.

Backlog Composition

Whether the backlog is aging in place or churning: the open-issue vs open-PR composition of the backlog per snapshot. A stable split with rising totals means items accumulate without being recycled; a shifting split means one queue is being drained into the other.

How to read: Each line is one backlog component's share of the combined issue+PR backlog; x = snapshot date, y = share. Example 1: A flat issue-share line near 98% means the backlog's composition is frozen — issues dominate and nothing structural is changing week to week. Example 2: A dropping PR-share line means pull requests are being merged or closed faster than new ones arrive, while the issue mountain stays.

Conclusion: The backlog is overwhelmingly issues (~98% of open items), and that split barely moves between snapshots: the org's backlog ages in place rather than churning. Combined with the drift panels above, the picture is a stable, slowly growing issue mountain tended by a small, diversified core — the org-shape context in which the correction labor of the Correction Anatomy page happens.

Back to overview