All-dictionary coverage
This tool extends the nine-chapter comparison to every CDSL v02 dictionary with a main source file. It keeps the MW §3 block vocabulary as the probe, but reports both structure and size: record count, entry length, block character mass, and entry-type inventory.
Trust Block
- Evidence:
src/data/dictionary-coverage.jsongenerated from local dictionary source files. - Limitations: coverage measures recoverable structural signals, not dictionary quality, authority, or complete semantic coverage.
- Validation: generated by
npm run build-coverage; checked bynpm testandnpm run build. - Owner repo:
csl-atlas. - Next use: inspect highlighted rows, then open exact dictionary source records before citing the pattern.
Corpus attestation by multiplicity (V4 strip)
How often the Digital Corpus of Sanskrit attests a union lemma, by how many dictionaries list it — dictionary-unique vocabulary is overwhelmingly corpus-invisible. Full analysis, per-dictionary shares, the Heritage witness cube, and the ghost-candidate queue: Ghost stock.
Union growth and Heaps saturation (V4 panel i)
Accumulating the 15 union dictionaries in publication order (dates from the
Lexicographic timeline), cumulative distinct lemmas follow a
saturating Heaps-type law — V(n) =
Trust Block
- Evidence:
src/data/lexico/heap_sat.json(PH8 HEAP-SAT packet) from the SanskritLexicography union headword backbone +data/dictionary_inventory.csvpublication years. - Limitations: the token axis counts union lemma listings, not corpus tokens; the specialised-break test is directionally positive but not significant (order-permutation p
, n=3 specialised dictionaries) — descriptive only; multi-volume spans collapse to one inventory year. - Validation: generated by
npm run build-heap-sat; checked bynpm run validate-heap-sat,npm test, andnpm run build. - Owner repo:
csl-atlas. - Next use: read the per-dictionary residuals before citing any "dictionary X added little" claim; the plateau answers "do I need more than MW + Apte?" for general dictionaries only.
Era-frequency signatures (V4 panel ii)
Which Sanskrit does each dictionary record? Each matched union lemma contributes its own normalized DCS-period share vector (type-weighted, so ubiquitous lemmas do not dominate); bars are the dictionary's mean profile, ticks the whole-union baseline. GRA and VEI shift hard toward Vedic; the indigenous kośas (SKD, VCP) toward the classical and late periods; the Petersburg line sits between. Match rates against the kosha frequency release are shown per dictionary — the unmatched stock is the corpus-invisible vocabulary characterised in Ghost stock.
Chronological centre of mass over the eight dated periods (bootstrap 95% CI; Epic/Classic are undated DCS layers and excluded from the score):
Trust Block
- Evidence:
src/data/lexico/period_signatures.json(PH3 FREQ-STRAT packet): SanskritLexicography union provenance × the frozen koshalemma_frequency.tsvDCS-period vectors, joined on the normalized SLP1 key. - Limitations: signatures describe each dictionary's corpus-visible slice (
% of union lemmas match a kosha row); DCS period sampling is uneven; family-level separation is descriptive only (Kruskal–Wallis H , p≈ , n=14 dictionaries) — the signal is per-dictionary canon (GRA), not language-pair family. - Validation: generated by
npm run build-period-signatures; checked bynpm run validate-period-signatures,npm test, andnpm run build. - Owner repo:
csl-atlas. - Next use: use the profiles as dictionary-chooser evidence ("GRA for the Ṛgveda, kośas for śāstra"), not as claims about each dictionary's own citations.
Fit and scale
Common-block population
Size by dictionary
Block size inside selected dictionary
Entry-type inventory
The current all-dictionary pass uses nine heuristic entry-type buckets: root/verbal lemma, three nominal genders, adjective, indeclinable/particle, compound/continuation, proper/encyclopedic, and other/untyped. This is a research scaffold: the counts identify where the type system transfers cleanly and where dictionary-specific review is needed.
Source: generated by npm run build-coverage from local ../csl-orig/v02 source files. Block character counts are overlapping signal measurements, not a partition of the entry. The fit scores are exploratory and should guide philological review, not replace it.