# KMGD v2 Methods

## 1. Purpose

KMGD v2 was designed to convert an earlier collection of downloadable mountain-reference tables into a source-traceable, versionable geographic dataset suitable for archival publication.

The main methodological objective is not to claim exhaustive coverage. It is to provide a curated reference set in which important numeric and classificatory statements can be traced to evidence and in which disagreement between sources is explicit.

## 2. Legacy migration

The legacy data contained separate “Top 30 highest peaks”, additional notable peaks, mountain ranges, selected glaciers, and general statistics.

During v2 migration:

- the “Top 30” ranking concept was removed because the table was not a strict national ranking;
- selected peak records were merged into one entity table;
- records outside Kyrgyzstan scope or with insufficiently established identity were removed from the core but retained in migration audit files;
- mountain ranges and non-range geographic areas were normalized into `mountain_units`;
- stable IDs were assigned before further revisions;
- later changes do not reuse IDs of removed legacy records.

## 3. Stable identifiers

Current prefixes:

- `KMGD-P####` — peaks/high points
- `KMGD-U####` — mountain units
- `KMGD-G####` — glaciers
- `KMGD-I####` — national indicators
- `KMGD-N####` — names
- `KMGD-S####` — sources
- `KMGD-V#####` — provenance records

Identifiers are intended to remain stable across future versions even when display names or canonical values change.

## 4. Entity model

### Peaks

A peak record represents one selected summit or a deliberately modeled high-point entity.

The peak table stores only the direct `mountain_unit_id`; mountain-unit names and mountain-system hierarchy are resolved through `kyrgyzstan_mountain_units.csv`. This avoids duplicate classification fields becoming inconsistent.

### Mountain units

`kyrgyzstan_mountain_units.csv` supports several unit types rather than pretending that every geographic entry is a range.

Current types are:

- `range`
- `massif`
- `climbing_region`

Parent-child links are represented through `parent_unit_id`.

### Glaciers

The glacier table contains selected well-documented large or scientifically monitored glaciers. It is not a national glacier inventory.

### National indicators

An indicator is included only when its definition and reference frame can be made explicit. Similar-looking statistics with different definitions are stored separately.

## 5. Canonical values

KMGD stores one canonical value where a usable editorial decision can be made.

A canonical value does **not** mean that every published source agrees. Published alternatives remain in `kyrgyzstan_provenance.csv`.

Examples of editorial treatment include:

- a map/reference elevation selected while expedition GPS readings are retained as alternatives;
- a corrected summit identity when a legacy value actually referred to a neighboring point;
- an official/institutional range maximum preferred over a secondary compilation;
- a glacier length represented as a range when the retained source itself gives a range.

No averaging is performed merely to reconcile disagreeing published values.

## 6. Evidence hierarchy

KMGD uses a project-specific evidence hierarchy.

### Tier A

Preferred high-authority evidence, including:

- official Kyrgyz government/statistical/encyclopedic sources;
- peer-reviewed scientific literature;
- scientific institutions and catalogues;
- intergovernmental sources when directly relevant.

### Tier B

Strong specialized evidence, including:

- American Alpine Journal and Alpine Journal;
- established alpine clubs and mountaineering institutions;
- detailed primary expedition/route reports;
- academic repositories and dissertations where directly relevant.

### Tier C

Useful specialized or regional references that may support cross-checking but are weaker as sole evidence for critical canonical values.

### Tier D

General compilations, encyclopedic web summaries, and peak databases used mainly for discovery or cross-checking.

The tiers are **editorial evidence categories for this dataset only**. They are not general rankings of websites or publications.

For the current pre-release peak/unit core, every canonical peak elevation and mountain-unit maximum has at least Tier A or B support.

## 7. Field-level provenance

Each evidence relationship may specify:

- entity type and ID;
- field name;
- source ID;
- value or statement reported by the source;
- source role;
- evidence relation;
- locator;
- derivation;
- editorial verification note.

The evidence relation is central to interpretation:

- `supports`
- `additional_support`
- `cross_check`
- `alternative_value`
- `conflicts`
- `derived_from_dataset`

A conflicting record is deliberately retained when it helps explain why a canonical value was chosen.

## 8. Naming

The v2 canonical naming policy favors a stable English display label that is supported by the retained evidence and useful for international data reuse.

`kyrgyzstan_names.csv` stores:

- canonical names;
- selected alternative names;
- transliteration variants;
- legacy dataset labels.

Legacy labels are preserved for migration traceability. They should not automatically be treated as exact synonyms: for example, an older broad massif/area label may have been replaced by a more precise range entity.

Kyrgyz- and Russian-language name coverage is not yet systematic and is therefore not filled speculatively.

## 9. Transboundary features

A feature can be important to Kyrgyz mountain geography without lying wholly inside Kyrgyzstan.

Peak relations currently distinguish:

- `within_kyrgyzstan`
- `border`
- `border_tripoint`

Mountain units distinguish:

- `within_kyrgyzstan`
- `transboundary`

For transboundary mountain units, the canonical maximum may describe the full unit rather than only the Kyrgyz portion. `elevation_scope`, `countries`, and provenance should be read together.

## 10. Glacier measurements

Glacier dimensions vary through time and depend on inventory method and outline date.

Therefore:

- dimensions are not filled merely to make the table visually complete;
- approximate values are flagged;
- ranges can be stored with `length_km_min` and `length_km_max`;
- area reference dates are retained when available;
- scientific monitoring relevance may be more important than a single current length for selected reference glaciers.

The national CAIAG inventory is represented separately in `kyrgyzstan_national_indicators.csv`.

## 11. Derived indicators

Derived values are marked with `derived=true`.

The formula is recorded in provenance where appropriate. Examples include:

- relative glacier-area change between two inventory baselines;
- modern glacier area as a percentage of national territory;
- number of represented peaks above 7,000 m.

A derived value should not be confused with an independently published statistic.

## 12. Validation

The pre-release workflow checks:

1. unique primary IDs;
2. peak → mountain-unit foreign keys;
3. mountain-unit parent foreign keys;
4. representative peak foreign keys;
5. glacier → mountain-unit foreign keys;
6. provenance → source foreign keys;
7. provenance → entity foreign keys;
8. names → entity/source foreign keys;
9. data-dictionary coverage of every released column;
10. presence of Tier A/B support for all canonical peak elevations and mountain-unit maxima;
11. preservation of documented alternative/conflicting values rather than overwriting them.

Validation reports are generated as separate CSV artifacts during release preparation.

## 13. Exclusions and scope control

A record may be excluded when:

- the summit/unit is outside Kyrgyzstan scope;
- the legacy entity cannot be independently established;
- the legacy label combines multiple distinct geographic objects;
- a source-backed definition cannot yet be made reproducible.

Exclusion from the core is preferable to retaining a visually complete but weakly defined record.

## 14. Versioning

KMGD v2 is intended for versioned archival publication.

Future versions may add coordinates, GeoJSON, more peaks, more glaciers, local-language names, or new evidence layers. Existing stable IDs should not be renumbered simply because display order or coverage changes.

For each public release, release files are frozen, checksummed, assigned citation/licence metadata, and validated as one package.


## Release-engineering rules

The public `kyrgyzstan_sources.csv` is restricted to sources that are actually referenced by the name or provenance layers.

An `alternative_value` evidence record must contain the alternative value in `reported_value`; an unspecific corroborating source is instead represented as `cross_check`.

For package interoperability, the public descriptor is `datapackage.json` with the `tabular-data-package` profile. Standard Table Schema foreign keys are encoded where the relationship is not polymorphic. Polymorphic links such as `provenance.entity_id` and `names.entity_id` remain documented and are validated by KMGD's release validator rather than expressed as a single Frictionless foreign key.

The standard filename `CITATION.cff` is used as an intentional exception to the otherwise consistent `kyrgyzstan_` filename prefix.
