L3 · Corpus & usage
Statistics over running-text usage: the DCS full corpus, SamudraManthanam's Sa↔Ru parallel corpus, frequency layers, and the still-open meter/accent statistics. Part of the statistics census overview (H817 WS1.3).
Statistics
Done
Partial
Not started
:::note
Trust block. Source: stats_census_register.csv,
rows where layer = L3, aggregated from VisualDCS, SamudraManthanam, and kosha.
n =
Headline magnitudes
How to read: log-scale bar of corpus-scale counts. Example: the 40.6M stop-word parallels dwarf the 5.7M token corpus itself by an order of magnitude — it is a pairwise/derived count, not raw text volume; read it as a derived-index size, not "more text than DCS holds."
Status breakdown
Full table
Download: stats_census_register.csv (full register, all layers) · Data downloads.
Open gaps
POS distribution per text moved from ○ to closed under H817 WS1.2 (now census-registered). Meter/prosody (SanskritKaraoke) and Vedic accent coverage (pending VedaWeb reuse) remain genuinely ○ not started — scheduled Q3 (WS3.5) per the roadmap, not blocked on anything this pass could resolve.
Chart density note: 2 Plot.plot calls (magnitude bar + status bar) —
justified per the same heterogeneous-units reasoning as the L1 page.