Orthographic drift explorer

The first atlas page about the gloss languages themselves: pre-reform German (Thier → Tier) and pre-1918 Russian (въ → в) spellings in the dictionary gloss text, counted against the frozen SanskritSpellCheck reform maps and read as a dating and descent signal (PH5 ORTHO-CLOCK, agenda §2/§3 V5).

It answers two testable questions — (i) the clock: does pre-reform density fall monotonically with publication date across PWG → PW → Nachträge → Schmidt? (ii) descent: do declared descendants carry elevated fossil orthography relative to date-matched independents, because copied German glosses carry the parent's spellings?

This is the meta-language layer — how the lexicographers spelled their German and Russian, not how they recorded Sanskrit. The Sanskrit-side house style lives on the Convention fingerprints page.

Verdicts — clock: · descent:

Drift vs year — the German census

Each point is one dictionary: dated-reform pre-1901/1996 spellings per 1,000 German gloss tokens, at the inventory mid-year of its publication span, with entry-level bootstrap 95% CIs (often narrower than the dot). The dashed line is the descriptive OLS fit (Spearman ρ , exact permutation p over all orderings).

The pattern is sharper than a date clock: the Böhtlingk lane stays uniformly fossil (PWG 14.8 → PW 17.5 → PWKVN 17.1 per 1k — his kürzere Fassung of 1879–89 is even more pre-reform than the 1855–75 original), while the non-Böhtlingk dictionaries modernise with date (GRA 10.4, CCS 6.5, SCH 4.8). Pre-reform density is a house-style clock, not a pure publication-date clock.

Era composition — each dictionary's orthographic epoch

Which reform regime dominates each dictionary's dated fossils. Every 19th-century dictionary is 1901-reform-dominated (th → t, c → k/z); Schmidt (1928) flips to 1996-ß-dominated with the 1901 signal collapsed — his epoch is readable from his own gloss text.

The census table

The descent contrast — refuted, and honestly bounded

With exactly one independent German dictionary in the corpus (GRA), no group-level permutation test is identifiable; the descent claim is tested as directional entry-level pair contrasts:

Cappeller declared CCS "nach den Petersburger Wörterbüchern bearbeitet", yet his 1887 German is less pre-reform than the independent GRA of 1873 — copying Petersburg content did not copy Petersburg spelling. Orthography follows the editor's own decade and house style, not the source text. The form-overlap column above says the same thing gently: even GRA (independent) shares 71% of its drifted forms with PWG, because the same high-frequency German words (Theil, thun, Thier) drift everywhere — fossil forms are shared German, not a descent fingerprint.

Top drifted forms — the searchable fossil list

The dated-reform fossils across all six dictionaries (old form · modern form · reform era · per-dictionary counts). Search matches old and modern forms. The era-unattributed corpus-mined map rows are excluded from this showcase (they remain in the all-map counts above).

Learner read. This table is why searching PWG for modern German fails: PWG never writes Teil, it writes Theil (1,245× across the roster). To search 19th-century German glosses, de-reform your query: tth inside Erbwörter, k/zc in Latinate words, ssß. The CSV doubles as a query-normalisation map.

Russian — Kossovich, the radical-reform case

The 1918 Russian reform abolished whole letters (ѣ, і, ѳ, ѵ, word-final ъ), so pre-reform Russian is detectable by definition, wordlist-free. Kossovich (1854) runs at per 1,000 tokens — roughly % of all Russian gloss tokens, an order of magnitude past the German maximum (/1k), because the hard-sign rule alone touches nearly every masculine noun.

Chart Trust Block


Generated by npm run build-ortho-drift. See docs/ATLAS_RESEARCH_AGENDA.md §2 PH5 / §3 V5 / §5d and docs/HYPOTHESIS_INDEX.md. CC-BY-SA-4.0.