All-dictionary coverage

This tool extends the nine-chapter comparison to every CDSL v02 dictionary with a main source file. It keeps the MW §3 block vocabulary as the probe, but reports both structure and size: record count, entry length, block character mass, and entry-type inventory.

Trust Block

Corpus attestation by multiplicity (V4 strip)

How often the Digital Corpus of Sanskrit attests a union lemma, by how many dictionaries list it — dictionary-unique vocabulary is overwhelmingly corpus-invisible. Full analysis, per-dictionary shares, the Heritage witness cube, and the ghost-candidate queue: Ghost stock.

Union growth and Heaps saturation (V4 panel i)

Accumulating the 15 union dictionaries in publication order (dates from the Lexicographic timeline), cumulative distinct lemmas follow a saturating Heaps-type law — V(n) = · n^ (log–log R² ). Points above the curve add more novel vocabulary than the recording process predicts: the largest breaks are the indigenous kośa SKD and the specialised BHS.

Trust Block

Era-frequency signatures (V4 panel ii)

Which Sanskrit does each dictionary record? Each matched union lemma contributes its own normalized DCS-period share vector (type-weighted, so ubiquitous lemmas do not dominate); bars are the dictionary's mean profile, ticks the whole-union baseline. GRA and VEI shift hard toward Vedic; the indigenous kośas (SKD, VCP) toward the classical and late periods; the Petersburg line sits between. Match rates against the kosha frequency release are shown per dictionary — the unmatched stock is the corpus-invisible vocabulary characterised in Ghost stock.

Chronological centre of mass over the eight dated periods (bootstrap 95% CI; Epic/Classic are undated DCS layers and excluded from the score):

Trust Block

Fit and scale

Common-block population

Size by dictionary

Block size inside selected dictionary

Entry-type inventory

The current all-dictionary pass uses nine heuristic entry-type buckets: root/verbal lemma, three nominal genders, adjective, indeclinable/particle, compound/continuation, proper/encyclopedic, and other/untyped. This is a research scaffold: the counts identify where the type system transfers cleanly and where dictionary-specific review is needed.


Source: generated by npm run build-coverage from local ../csl-orig/v02 source files. Block character counts are overlapping signal measurements, not a partition of the entry. The fit scores are exploratory and should guide philological review, not replace it.