Orthographic drift explorer
The first atlas page about the gloss languages themselves: pre-reform German (Thier → Tier) and pre-1918 Russian (въ → в) spellings in the dictionary gloss text, counted against the frozen SanskritSpellCheck reform maps and read as a dating and descent signal (PH5 ORTHO-CLOCK, agenda §2/§3 V5).
It answers two testable questions — (i) the clock: does pre-reform density fall monotonically with publication date across PWG → PW → Nachträge → Schmidt? (ii) descent: do declared descendants carry elevated fossil orthography relative to date-matched independents, because copied German glosses carry the parent's spellings?
This is the meta-language layer — how the lexicographers spelled their German and Russian, not how they recorded Sanskrit. The Sanskrit-side house style lives on the Convention fingerprints page.
Verdicts — clock: · descent:
Drift vs year — the German census
Each point is one dictionary: dated-reform pre-1901/1996 spellings per 1,000
German gloss tokens, at the inventory mid-year of its publication span, with
entry-level bootstrap 95% CIs (often narrower than the dot). The dashed line is
the descriptive OLS fit (Spearman ρ
The pattern is sharper than a date clock: the Böhtlingk lane stays uniformly fossil (PWG 14.8 → PW 17.5 → PWKVN 17.1 per 1k — his kürzere Fassung of 1879–89 is even more pre-reform than the 1855–75 original), while the non-Böhtlingk dictionaries modernise with date (GRA 10.4, CCS 6.5, SCH 4.8). Pre-reform density is a house-style clock, not a pure publication-date clock.
Era composition — each dictionary's orthographic epoch
Which reform regime dominates each dictionary's dated fossils. Every 19th-century dictionary is 1901-reform-dominated (th → t, c → k/z); Schmidt (1928) flips to 1996-ß-dominated with the 1901 signal collapsed — his epoch is readable from his own gloss text.
The census table
The descent contrast — refuted, and honestly bounded
With exactly one independent German dictionary in the corpus (GRA), no group-level permutation test is identifiable; the descent claim is tested as directional entry-level pair contrasts:
Cappeller declared CCS "nach den Petersburger Wörterbüchern bearbeitet", yet his 1887 German is less pre-reform than the independent GRA of 1873 — copying Petersburg content did not copy Petersburg spelling. Orthography follows the editor's own decade and house style, not the source text. The form-overlap column above says the same thing gently: even GRA (independent) shares 71% of its drifted forms with PWG, because the same high-frequency German words (Theil, thun, Thier) drift everywhere — fossil forms are shared German, not a descent fingerprint.
Top drifted forms — the searchable fossil list
The dated-reform fossils across all six dictionaries (old form · modern form · reform era · per-dictionary counts). Search matches old and modern forms. The era-unattributed corpus-mined map rows are excluded from this showcase (they remain in the all-map counts above).
Learner read. This table is why searching PWG for modern German fails: PWG never writes Teil, it writes Theil (1,245× across the roster). To search 19th-century German glosses, de-reform your query: t → th inside Erbwörter, k/z → c in Latinate words, ss → ß. The CSV doubles as a query-normalisation map.
Russian — Kossovich, the radical-reform case
The 1918 Russian reform abolished whole letters (ѣ, і, ѳ, ѵ, word-final ъ), so
pre-reform Russian is detectable by definition, wordlist-free. Kossovich
(1854) runs at
Chart Trust Block
- Claim: pre-reform gloss-language spelling density and era composition, per
dictionary, as a dating/descent signal — clock
, descent ( derived). - Evidence label:
derived— deterministic map-membership census over csl-orig gloss text; no model inference on this page. - Source files:
de_reform_map.tsv(pairs) · ru_reform_map.tsv(pairs) — owned by SanskritSpellCheck (A37 lane), consumed frozen; csl-orig v02 gloss text (pwg, gra, pw, pwkvn, ccs, sch); SamudraManthanam kossovich.jsonl. - Generated by:
npm run build-ortho-drift→data/lexico/ortho_drift.json(+.source.jsonprovenance envelope with sibling commits). - Statistics: entry-level bootstrap 95% CIs (B=1,000, fixed-seed mulberry32); exhaustive n!-permutation Spearman for the tiny-n date regression; directional entry-level permutation pair tests (B=1,000) — Dror et al. 2018 protocol, Bollmann 2019 framing for historical normalisation.
- Validation:
npm run validate-ortho-drift(rate/share/era-sum coherence, CI bracketing, residuals-vs-fit, verdict-vs-statistics consistency, live map row-count cross-check when the sibling checkout is present);npm test;npm run build. - Known limits: the meta-language axis dates the editing, not the Sanskrit; map-membership counting undercounts unmapped inflected drift; the era-unattributed map rows are excluded from the headline clock. Full list below.
- Review status: machine-reviewed; the lineage labels (progenitor / descendant / independent) are documented genealogy (title pages, L0 layer), not inferred here.
- Owner repo:
csl-atlas; reform maps stay upstream in SanskritSpellCheck — the atlas owns only this cross-dictionary census layer. - Next action: a second independent German dictionary (outside CDSL) would make the descent contrast group-testable; KNA (1893) text would give Russian a second dated point.
- Boundary note: gloss-language orthography only — Sanskrit-side drift is VisualDCS territory; csl-orig is never written.
Generated by npm run build-ortho-drift. See
docs/ATLAS_RESEARCH_AGENDA.md
§2 PH5 / §3 V5 / §5d and
docs/HYPOTHESIS_INDEX.md.
CC-BY-SA-4.0.