Created: 25-09-2026 · Last updated: 25-09-2026
title: Access structures — printed sort order toc: false
Access structures — what rule produced the printed order?
Wiegand's Zugriffsstruktur is the ordering a reader relies on to find an entry.
This page carries the census behind
docs/ACCESS_STRUCTURES_SORT_ORDER_CENSUS_24-09-2026.md:
for MW, PWG, PW (PWK), AP, WIL, SKD and VCP it recovers the headword order actually
printed (the csl-orig record order), tests it against a grid of 72 varṇamālā rules plus
two controls, and reports the best-fit rule, its violation rate, the evidence that
separates it from its rivals, and every counterexample.
The two versified kośas (ARMH, ABCH) are ordered by concept, not by alphabet, and are measured elsewhere. Sanskrit is displayed in IAST throughout; the SLP1 sort key is kept as a muted secondary column.
Trust Block
- Evidence:
data/lexico/access_structures.jsonand the two CSV siblings, generated byscripts/lexico/m11_access_structures.py(npm run build-access-structures, stdlib only, deterministic) straight from csl-orig. - Limitations: printed order is assumed to be csl-orig record order; running heads are not in the data; factors marked
undeterminedhave too few disagreeing adjacent pairs to decide. - Validation: hand-derived pins in
tests/forensic/test_m11_access_structures.py. - Owner repo:
csl-atlas(handoffs H5331, H5409). - Next use: read a rate as how well one rule fits, never as a quality score for the dictionary — AP's 9.4 % is a nested macrostructure, not sloppy alphabetisation.
At a glance
Best fit per dictionary
A violation is an adjacent pair of headwords that goes down under the rule, counted inside one alphabet run. Displaced is the minimum share of headwords that would have to move for the order to satisfy the rule. The roman control is the same order read as IAST in Latin-alphabet order: at 31–37 % everywhere, it is what shows that every one of these dictionaries follows the varṇamālā.
Factor status — what the order can and cannot decide
Overall rates cannot settle a factor: most adjacent pairs never touch ṃ, ḥ or a geminate
after r. Each factor is judged on pairwise contrasts — only the pairs where two of its
policies disagree, the other factors held at best fit. Fewer than
undetermined; a winner at 3 : 1
or better is clear. partial means some alternatives are rejected and others are
indistinguishable from the chosen policy.
AP's root nests — and why vowel grades do not explain them
Apte prints derivatives under their verbal root, in derivational rather than alphabetical
order, so its rate is a macrostructure effect. H5409 tested whether the residue is
derivatives carrying the root vowel in guṇa or vṛddhi grade (√kṛ nesting kāra,
karaṇa): the heads_nest view collapses those into the root's position. It removes only
an eighth of the disorder, so refutation condition 3 of the methods page fired —
the vowel-grade explanation is refuted, and prefix-family grouping is the untested
alternative.
Counterexample browser
Every within-run descent of the best-fit rule in each dictionary's primary view, with its
L-id, page reference and the pair in IAST. kind is block when the pair sits inside a
stretch of ≥ isolated otherwise.
sensitive_to_anusvara_visarga_policy marks the pairs that some other rule in the grid
would order correctly — the pairs that carry the evidence about ṃ and ḥ.
The full rate table
Every dictionary × view × rule combination the census scored — the raw material behind the best-fit choices above.
Method in one paragraph
The sort key is k1, the accentless normalised SLP1 headword; consecutive records sharing
a key collapse into one position. Alias records ({{Lbody=N}}, spellings added by the
digitisers) are dropped from every view but all; MW and AP additionally give every run-on
its own record, so their primary view keeps level-1 heads only, and AP's drops paragraph
run-ons and root-nest derivatives on top of that. A supplement or Nachträge run is detected
as an alphabet restart — the initial letter goes backwards and the next
Dr. Mārcis Gasūns