Created: 24-09-2026 · Last updated: 24-09-2026


title: Methods — the kośa as a macrostructural type

The versified synonymic kośa as a macrostructural type

The Amarakośa, Halāyudha's Abhidhānaratnamālā (ARMH) and Hemacandra's Abhidhānacintāmaṇi (ABCH) are dictionaries whose meaning lives in their arrangement: a concept-ordered hierarchy of books (kāṇḍa), sections (varga), verses and synonym-sets, with liṅga (gender) marking riding on the words. The atlas measured that arrangement in the A06 kośa macrostructure paper; this page states it as a model — a schema every kośa instance must satisfy — and separates what the sources show from what we infer.

Trust Block

1. Two macrostructural types

A European dictionary is semasiological: it starts from the word and has one ordering device, the alphabetised headword; the lexicographic work is done in the entry (senses, glosses, citations). A kośa is onomasiological: it starts from the concept, and the lexicographic work is done by placing a word — which book, which section, which verse, next to which synonyms, in which gender.

Level Kośa (onomasiological, versified) European (semasiological, alphabetical)
Top division kāṇḍa — region of the universe (Amara, Halāyudha) or hierarchy of beings (Hemacandra) letter of the alphabet
Section varga — a concept field (heaven, time, the body, …); upavarga in Hemacandra's animal book none (letter ranges)
Unit of text verse (śloka), numbered within its section entry (article)
Unit of meaning synonym-set: the names of one concept sense within an entry
Homonymy a dedicated nānārtha section, ordered by the word's final sound homonym numbers or senses inside one entry
Grammar liṅga stated by form, association or explicit word (striyām, klībe) part-of-speech label in the entry
Findability memorised verse; no alphabetical access alphabetical lookup

The schema encodes the left column as kāṇḍa → varga → (upavarga) → verse-group → set → member, with four set kinds: synonym-set, homonym-sense (one headword + a gloss naming one meaning), indeclinable-set, and unsegmented-verse for a digitization that does not mark set boundaries (ARMH). Each instance also carries an orderingDevices table in which every device is labelled observed, inferred or absent and assigned to a layer: the source text, the digitization markup, or our measurement.

2. The ordering devices, observed versus inferred

Observed in the source text — present in the words of the kośa itself:

  1. the kāṇḍa and varga divisions. Section headings name all 24 vargas of the Amarakośa; 23 open with atha … vargaḥ (once in sandhi, athāvyayavargaḥ) and 23 close with iti … vargaḥ (the bhūmi-varga has no opening colophon, the avyaya-varga no closing one);
  2. the verse and its number. Amara's numbers restart at every one of the 23 varga boundaries and, in this digitization, run in steps of exactly one inside each varga: 1,432 full verses inside the verse-groups, 1,444 in the file, the other 12 standing outside any verse-group (the preface, for example). Hemacandra's numbers, by contrast, run on through all 14 section boundaries of ABCH without restarting;
  3. the homonym section (nānārtha-varga), whose opening verse announces the arrangement by final sound (kāntādi);
  4. the indeclinable section (avyaya-varga), the last varga of this digitization — the text itself (kāṇḍa 3, v. 1) names one more, the liṅgādisaṅgraha-varga, after it;
  5. the rule of gender marking: Amara's paribhāṣā (vv. 3–5) says liṅga is known by form, by association with a neighbouring word, or by an explicit statement.

Observed in the digitization markup — explicit in the files, but supplied by the annotators, not by the author:

  1. the segmentation of a verse into synonym-sets (<eid>): 5,590 sets in 2,359 verse-groups for the Amarakośa;
  2. the per-word liṅga tag (puM, strI, klI, tri, a, with dvi/ba for number): 14 distinct tags, all parsed;
  3. Hemacandra's upavarga tier in the animal book (e.g. pañcendriyasthalacara).

Inferred by measurement — patterns we establish; the text states none of them, or (for the homonym section) only the principle:

  1. No alphabetical device. Adjacent synonyms are in alphabetical order no more often than chance.
  2. How strictly the a-tergo order of the homonym section holds. The kāntādi verse announces the principle. That it holds almost perfectly, that it runs in two series, and that its exceptions are the traditional letter equivalences is what we measure.
  3. Gender contiguity in the Amarakośa: words of one gender stand together.

Absent: an upavarga tier in the Amarakośa, the liṅgādisaṅgraha-varga in this digitization, and, for ARMH, varga divisions, synonym-set boundaries and gender tags.

3. Measurements

3.1 No alphabetical device

Share of adjacent words whose Sanskrit (varṇa) collation does not decrease. An alphabetical dictionary sits near 1; an unordered list near 0.5.

MW scores 0.939 (not 1.0, because compounds are nested under their first member); the three kośas score 0.49–0.50. Alphabetical order is not a kośa device, at any level.

3.2 The homonym section is ordered by the final sound

Over the 861 homonym headwords (1,995 senses), the final consonant is in varṇa order on 0.991 of adjacent steps; the initial letter only on 0.536, which is chance. The section is sorted a tergo. It is also two series, not one: 802 nominal headwords run from nāka (final -k-) to the -h- finals, then a second series of 59, 98% indeclinable, starts again from āṅ, āḥ, ku, dhik, ca …. In the first series only five steps go backwards, and four of them fall under the equivalences the grammatical tradition itself allows — ḍ = l (ilā → kṣveḍā) and b = v (gandharva → kambu, pūrva → kumbha, sattva → klība). One (jihma → uṣṇa) is unexplained.

The order is measured; that it was intended is supported by the section's own opening verse, and its exceptions pattern with a known convention, but the reading of the two series as a deliberate "nominal, then indeclinable" design is an inference.

3.3 Gender contiguity

In a multi-gender synonym-set, does each gender form one unbroken run? The chance baseline is exact: for a set of n words in k genders with counts c₁ … cₖ, a random order keeps every gender together with probability k! · Π cᵢ! / n!.

In the Amarakośa 0.90 of 756 multi-gender sets keep each gender together, against 0.68 by chance. This is how the verse carries gender: a run of masculines, then dve striyām ("the two [are] feminine"), then klībe ("in the neuter"). In Hemacandra the effect nearly disappears (0.67 against 0.62). Why is not tested here; one candidate is that Hemacandra treats gender in his separate Liṅgānuśāsana, so his verse need not carry it. Masculine-first is at chance in both (0.36 vs 0.35; 0.34 vs 0.34): the device is grouping, not a fixed gender order.

3.4 The model instance, one set of each kind

4. What the model does not claim

  1. The sets are editorial. A verse such as svar avyayaṃ svarga-nāka-tridiva-… does not mark where one concept ends and the next begins; the <eid> segmentation is the annotators' reading (Shivja S. Nair's database for the Amarakośa, the CDSL markup for ABCH). The model records it as digitization-markup, not source-text.
  2. The gender tags are editorial too. The text supplies gender by the three means of its paribhāṣā; the tag on each word is an annotator's resolution of them. The contiguity result (§3.3) is therefore a property of text plus annotation.
  3. Section types are read from labels. nānārtha → homonymic and avyaya → indeclinable are read from the varga names; ARMH's fifth kāṇḍa is typed homonymic from its …api… verse formula (A06 §4.2), so its basis is content, not section-label.
  4. Counts are per digitization. A "set" in AMAR/ABCH and a "verse" in ARMH are different units (A06 §4.3); the schema keeps them apart through digitizationModel and the unsegmented-verse kind.

5. Reproduce

The build needs the sibling checkouts ../AMAR (sanskrit-kosha data) and ../csl-orig; the validator and tests need only the committed JSON.

python scripts/lexico/m10_kosa_macrostructure_model.py
npm run validate-kosa-model
node --test test/kosa-model.test.mjs

Dr. Mārcis Gasūns