L1 · Lexicon text
Statistics over the 44 digitized dictionaries: headwords, overlap, markup, citations, corrections, and the sense/definition-level gaps still open. Part of the statistics census overview (H817 WS1.3).
Statistics
Done
Partial
Not started
:::note
Trust block. Source: stats_census_register.csv,
rows where layer = L1, aggregated from SanskritLexicography, csl-atlas, csl-orig,
CORRECTIONS, and csl-observatory itself. n =
Headline magnitudes
How to read: log-scale bar of every L1 statistic that has a plain numeric count (citations, headwords, tag hits, n-grams). Non-numeric rows (status-only or percentage rows) are omitted here and appear in the full table below. Example: the markup-tag census (17.5M tag hits) and the n-gram oracle (6.66M n-grams) dwarf the union-headword count (323k) by more than an order of magnitude — expected, since one headword carries many markup hits.
Status breakdown
Full table
Download: stats_census_register.csv (full register, all layers) · Data downloads.
Open gaps
The two ○ not-started rows (definition typology, per-dict editorial fingerprint) are Q2/analytical-layer work per the roadmap, not Q1 — see Part II Q2. Sense/polysemy per dict is genuinely capped at 11/44 dicts — the remaining 33 have no structural sense-marking to count against (dead end recorded in the roadmap's H817 WS1.2 pass, not a to-do).
Chart density note: 2 Plot.plot calls — a magnitude bar (comparable
numeric rows only) + a status bar. The remaining rows are non-numeric
(percentages, qualitative status) and are covered by the full table, not
manufactured into additional bars.