Ghost stock — is dictionary-unique vocabulary real?
42% of the cross-dictionary union's lemmas appear in exactly one dictionary. Are they recorded language, or lexicographic sediment — inherited citation-forms, artefacts, ghost words? This page joins the union headword backbone against two independent modern witnesses: the Digital Corpus of Sanskrit attestation flags (via the frozen VisualDCS summary) and, for MW, the Sanskrit Heritage lexicon crosswalk (agenda hypotheses PH4 GHOST-STOCK + PH6 HERITAGE-WIT).
Of
union lemmas, ( %) are DCS-attested — but the rate climbs monotonically from % at n_dicts = 1 to 100% for lemmas shared by 13+ dictionaries. Dictionary-unique vocabulary is overwhelmingly corpus-invisible.
Attestation by multiplicity (PH4)
Share of lemmas DCS attests, per number of dictionaries listing the lemma. Whiskers are Wilson 95% score intervals (mostly narrower than the dots).
Which dictionaries' unique stock is real? (per-dict)
Among each dictionary's unique lemmas (n_dicts = 1), the share DCS attests. Note the honest surprise: the highest unique-and-attested shares belong to MW and MD (general dictionaries), not to the specialised lexica the agenda predicted — though DCS under-samples Buddhist and technical literature, which caps what BHS-style uniqueness can show (see Trust Block limitations).
Logistic model
Descriptive logistic regression of attested on n_dicts plus
family-presence indicators (deterministic IRLS; McFadden pseudo-R²
The Heritage triangulation (PH6) — MW lemmas only
For the
Ranked ghost-candidate queue — inferred, not asserted
The triple filter MW-unique ∧ Heritage-uncovered ∧ DCS-unattested isolates
explicit (crosswalk-missing (
Chart Trust Block
- Claim: dictionary-unique headwords are disproportionately corpus-unattested (PH4), and Heritage non-coverage triangulates the same ghost stratum for MW, yielding a ranked candidate queue (PH6).
- Evidence label:
derivedfor the attestation rates, strata, cube, and model (deterministic joins + arithmetic on committed inputs);inferredfor the ghost-candidate queue — no row is asserted as a ghost word until a human source read (H5 discipline) decides it. - Source files:
union_headwords.tsv(normalized lemmas; owned by SanskritLexicography), mw_heritage_crosswalk.tsv(same owner),data/dcs/dcs_lemma_summary.json(lemmas, frozen VisualDCS export); committed packet src/data/ghost-stock/ghost_stock.jsonwith provenance in its.source.jsonsidecar. - Generated by:
npm run build-ghost-stock - Validation:
npm run validate-ghost-stock— strata/cube/queue arithmetic, Wilson-CI bracketing, logistic convergence, plus a full rebuild cross-check against the sibling checkout when present; the logistic fit is independently reproduced (NumPy IRLS) to 4 decimals. - Known false positives: none at the join level — attestation and coverage are read directly off the committed witness files, not inferred.
- Known false negatives: DCS samples the transmitted literature — "unattested" means absent from that corpus release, not absent from Sanskrit; Buddhist/technical vocabulary is under-sampled, which is exactly why BHS's unique stock scores low here. Homonyms collapse onto one normalized key (attested if any homonym is). Heritage coverage is MW-keyed and mirror-derived, not a live INRIA query.
- Review status: machine-reviewed (deterministic validator); the ghost-candidate queue awaits human spot-check — route via the H5 anomaly review discipline before any correction or paper claim.
- Owner repo:
csl-atlas(join + rendering); the union backbone and crosswalk are owned bySanskritLexicography; corpus attestation belongs to VisualDCS per the consumption contract. - Next action:
/spot-check-samplea stratified sample of theexplicittier against MW source lines; then fold verdicts into the H5 queue. - External dependencies:
SanskritLexicography(union + crosswalk), VisualDCS summary (in-repo frozen copy). - Boundary note: no DCS ingestion here (VisualDCS contract); Heritage is cited and joined, never cloned (LGPLLR).