CSL Observatory 13 years of Cologne Digital Sanskrit Lexicon

POS distribution by text

UD part-of-speech share across the 270 texts of the Digital Corpus of Sanskrit (DCS). Census magnitude bars only answered “how many tokens”; this page answers which texts skew noun-heavy vs verb-heavy — the register/genre signal behind the L3 POS feed.

Texts

UPOS rows

Tokens

UPOS tags

:::note Trust block. Source: data/pos_distribution_per_text.tsv (loader: observatory/site/src/data/pos_distribution_per_text.csv.py, read-only). Generator: scripts/pos_distribution_per_text.py from VisualDCS dcs_full.sqlite. n = rows · texts · tokens. Data date: 13-07-2026 (H817 WS1.2). DCS tags unaccented text — coarse 11-way UD upos only; no class I/VI or aorist/perfect split. :::

Top texts — UPOS stack

Stacked shares for the largest texts by token count (default top 25). Epic narrative tends toward higher VERB%; medical/śāstra denser in NOUN%.

How to read: Each horizontal bar is one text; segments are UPOS shares of that text’s tokens. Example 1: A long NOUN segment with a short VERB segment is a catalogue-like or technical register. Example 2: VERB near ~20% with NOUN ~40% matches the epic average (Mahābhārata, Rāmāyaṇa).

Conclusion: The stack separates catalogue/śāstra (noun-dominant medical and purāṇa blocks) from narrative epic and kathā (higher VERB%). Register, not dictionary markup, is what moves the POS profile.

Heatmap — text × UPOS (top 40)

Top 40 texts by total tokens; remaining texts are omitted so the grid stays readable.

How to read: Row = text, column = UPOS, colour = pct_of_text. Example 1: A bright NOUN cell on a medical text is high nominal density. Example 2: A bright PART/PRON cell marks formulaic particle-rich prose or dialogue.

UPOS share distribution across texts

Each point is one text’s pct_of_text for that UPOS — the spread of genre, not the corpus mean alone.

How to read: Box (or jittered dots) of percentage by UPOS. Example 1: A tight NOUN box around 40% means most texts sit near the epic mean. Example 2: A long upper whisker on NOUN is a technical text pulling far above the median.

Outlier ranks — NOUN% and VERB%

Texts ranked by noun share and verb share (among texts with ≥500 tokens so tiny fragments do not dominate).

How to read: Horizontal bars of share for the top/bottom tails. Example 1: Highest NOUN% texts are typically medical or list-like. Example 2: Highest VERB% texts lean narrative or ritual instruction.

Conclusion: The same feed that looked flat at census level (NOUN ~42%, VERB ~18% corpus-wide) is text-heterogeneous: medical and technical works pull NOUN% into the 50s; narrative and some ritual texts push VERB% above the epic band. Genre skew is the analytical payload of this page.

Data table

Download source TSV: pos_distribution_per_text.tsv · report: reports/pos_distribution_per_text.md · Data downloads · sibling census: L3 Corpus.