Citation canon explorer

The <ls> citation graph across eleven Cologne Sanskrit dictionaries, read as a single dictionary × text matrix: which canonical text each dictionary cites, how often, and how the whole apparatus is shaped. This is the first-class view of the data/citations/ graph behind the A50 paper.

It answers one testable question — PH1 CANON-CORE: is the shared canon a nested core–periphery (each dictionary's cited texts approximately a subset of the next-broader one's — one canon in additive strata), or is it modular (dictionaries carry partly disjoint tradition communities)?

This is the text graph — what the dictionaries quote. It is a different object from the Citation apparatus page, which measures apparatus style (density, breadth, siglum overlap). Read the two together.

The matrix is dictionaries × texts with citation edges (fill %). Dictionary codes are the <ls>-tagged Cologne sources; text names are already IAST.

Topology verdict —

The binarised matrix is scored two ways and each is compared to degree-preserving (fixed-fixed) permutation nulls that hold every dictionary's breadth and every text's popularity fixed:

Canon heatmap — dictionaries × most-cited texts

Rows (dictionaries) are ordered by apparatus breadth, top to bottom; columns (texts) by how many dictionaries cite them, left to right — the packing that makes a nested matrix look like a staircase and a modular matrix look blocky. Colour is log citation count.

Empty cells are true zeros (that dictionary does not cite that text in its tagged apparatus). A staircase of decreasing fill down and to the right would mean nesting; discrete blocks would mean disjoint communities.

Canon curve — how widely shared is each text?

The number of texts cited by exactly k of the dictionaries. A fat left tail (few universal texts, most texts private to one dictionary) is the signature of a thin shared head over idiosyncratic tails.

Per-dictionary fingerprint

Breadth (distinct texts cited), total citation volume, coverage of the shared core (the most widely-cited texts), and each dictionary's five heaviest sources.

Tradition communities — naming the modular split

The topology verdict above says the apparatus is modular: the dictionaries carry partly disjoint tradition communities. This panel names them, joining a curated text → tradition overlay onto the citation edges and reading off, per dictionary, how its citation volume splits across traditions.

⚠️ This map is ${trad.evidenceLabel} (). The ${trad.reviewState.taggedTexts} text→tradition assignments are scholarly proposals routed to human review (agenda backlog #9); unreviewed tags are shown as inferred, never asserted as fact. Confidence: high · medium · low. Shares are over the tagged texts (the modular signal — shared head + each dictionary's heaviest sources), not the full -text graph.

Each dictionary's bar is its citation profile across traditions — bhs reads as an almost pure Buddhist community, the Apte pair (ap/ap90) and lrv as classical-kāvya, mw/md as Vedic (over a small tagged sample — see the limitations below). Distinct blocks, not one shared ladder: the modular verdict made concrete.

The tagged texts

The ${trad.reviewState.taggedTexts} curated assignments behind the panel, with proposed tradition, confidence, and review state. This is the map A50 §4 cites; it is regenerated from data/citations/tradition_tags.tsv by npm run build-tradition-tags.

Most-cited texts — the shared reading list

The top canonical texts by number of dictionaries citing them, then by in-graph citation volume. The head of this table is the reading a learner can trust every dictionary to support.

Chart Trust Block


Generated by npm run build-citation-canon. See docs/ATLAS_RESEARCH_AGENDA.md §2 PH1 / §3 V1 and docs/articles/A50_ls_citation_frequency_graph.md. CC-BY-SA-4.0.