Etymology style — Nirukta markers vs Western cf.
A frozen exploratory witness of how five Cologne dictionaries mark
etymology / derivation. The claim is descriptive, not a new inference:
WIL (1832) still writes Nirukta-style affix notation inside .E. blocks
(aff., c., neg. …), while later Western editions (MW72, MW) shift to
comparativist cf. cross-references. Apte editions sit in between with sparse
morphology tags and almost no cf..
Source probes live in
csl-observatory as frozen
CSVs; this atlas page only renders them (dictionary microstructure belongs
here per
BOUNDARY_RULES.md).
Not a full 44-dict census.
Nirukta-style .E. share by dictionary
Percent of entries carrying a Nirukta-style etymology block (Wilson's .E.
notation). Only WIL is non-zero in this spike — the later sample has already
dropped the indigenous affix apparatus.
Western cf. count by dictionary
Raw cf. hit counts (comparativist cross-reference style). MW72 and MW dominate;
WIL has zero in this probe.
Timeline — year × Nirukta .E. share
Connected scatter (slope chart) of the same five dictionaries by publication year. The drop from WIL 1832 (~89%) to every later point (0%) is the visual of the style shift; it is not a regression over the full CDSL set.
WIL Nirukta tokens (top abbreviations)
Top tokens counted inside WIL .E. blocks. aff. (affix / pratyaya) alone
is ~14.7k — the indigenous derivation apparatus is affix-notation-heavy, not
cf.-heavy.
Full tables
Marker summary (5 dictionaries)
WIL token inventory
Chart Trust Block
- Claim: among a five-dictionary exploratory sample, WIL (1832) is the only
dictionary with a high share of Nirukta-style
.E.etymology blocks (~88.9%), while MW72/MW show largecf.counts and zero Nirukta-block share; WIL's top.E.token isaff.(affix). - Evidence label:
derived— deterministic token/marker counts from frozen observatory probes; witness grade = exploratory spike (not a full census, not a model). - Source files: vendored copies under
src/data/witness/with provenance headers pointing atcsl-observatory/data/etymology_marker_preliminary.csv(commit582f5337) andcsl-observatory/data/wil_nirukta_tokens.csv(commit8c4b78be). n = 5 dictionaries / 16 WIL tokens. - Generated by: no atlas builder — FileAttachment of frozen CSVs only (H1525).
- Validation:
npm run buildmust include this page; CSV provenance headers present; plot count ≥ 4. - Known false positives: a
cf.string match can hit non-etymological uses; morphology tags (caus/pass/desid) are co-present columns, not a style classification. - Known false negatives: only five dictionaries were probed — the other ~39 CDSL dictionaries are invisible here; orthographic reform maps (agenda V5) are out of scope.
- Review status: machine-rendered exploratory witness; not human-reviewed as a full etymology typology.
- Owner repo:
csl-atlas(rendering); probes owned bycsl-observatory. - Next action: if a paper needs the style claim, extend the probe to a pre-registered dict set (or re-run a documented extractor) before promoting beyond "exploratory spike"; do not re-extract in-page.
- External dependencies: csl-observatory frozen CSVs only (no live scrape).
- Boundary note: dictionary microstructure markers belong in the atlas;
org metrics stay in the observatory
(
DICTIONARY_STRUCTURE_MOVED.md).
Related
- WIL dictionary chapter — Wilson as the base of the European line
- Dictionary genealogy — content inheritance, not etymology style
- Convention fingerprints — house-style cladogram (Patel L0)
- Structural register — citation × grammar family scatter
- Lemma dossier — look up a headword across dictionaries