Chapter 7: Data Collection Specifications

ICO Std 2002:2026 — Chapter 7


This chapter specifies the requirements for data collection, processing, quality assurance, ethical compliance, and archival in ranking systems. It establishes a comprehensive data governance framework that ensures the integrity and trustworthiness of the data underpinning ranking results.


7.1 Data Source Classification

All data sources used in a ranking system shall be classified according to the three-tier system defined below. The classification determines the evidentiary weight and verification requirements applicable to each source.

7.1.1 Tier 1 Data Sources

Tier 1 data sources are the most authoritative and reliable sources available.

Tier 1 data sources include:

a) Official statistics: Data published by national statistical offices, intergovernmental organisations, and regulatory bodies in their official capacity.

b) Audited financial data: Financial statements and reports that have been independently audited in accordance with recognised auditing standards.

c) Administrative records: Data generated by governmental or institutional administrative processes, such as registration records, licensing data, and tax filings.

d) Peer-reviewed research: Data published in peer-reviewed academic journals or reports, where the underlying methodology has been subject to independent scrutiny.

7.1.2 Tier 2 Data Sources

Tier 2 data sources provide valuable information that may not have the same level of institutional authority as Tier 1 sources.

Tier 2 data sources include:

a) Public databases: Databases maintained by research institutions, non-governmental organisations, or industry associations that are publicly accessible.

b) Academic literature: Published academic works that have not undergone formal peer review for the specific data used, or where the data is secondary.

c) Industry reports: Reports produced by industry analysts, consultancies, or trade associations.

d) Self-reported data with independent verification: Data provided by ranked entities where the ranking entity has conducted independent verification of a representative sample.

7.1.3 Tier 3 Data Sources

Tier 3 data sources are the least authoritative and shall be used only when Tier 1 and Tier 2 sources are insufficient.

Tier 3 data sources include:

a) Media reports: Data derived from news articles, press releases, or other media publications.

b) Social media data: Data derived from social media platforms, online forums, or user-generated content platforms.

c) Unverified self-reported data: Data provided by ranked entities without independent verification.

d) Expert estimates: Quantitative estimates provided by individual experts in the absence of empirical data.

7.1.4 Evidentiary Rules for Data Source Tiers

The following rules govern the use of data from different tiers:

a) Tier 1 data shall be used as the primary source wherever available.

b) Tier 2 data may be used as a supplementary or alternative source when Tier 1 data is unavailable or insufficient.

c) Tier 3 data may only be used when:

1) no Tier 1 or Tier 2 data exists for the indicator; and

2) the use of Tier 3 data is necessary for the ranking purpose; and

3) the limitations of Tier 3 data are disclosed in the ranking publication.

d) When data from multiple tiers is available for the same indicator, the ranking entity shall document the rationale for the tier selected.

e) The proportion of data points sourced from each tier shall be reported in the methodology documentation.


7.2 Data Collection Process

Data collection shall follow a structured process that ensures completeness, accuracy, and traceability.

7.2.1 Collection Plan Development

Prior to data collection, the ranking entity shall develop and document a data collection plan that includes:

a) the list of indicators and the data requirements for each indicator, including:

1) the data elements to be collected;

2) the data source tier and specific source for each element;

3) the expected data format and unit of measurement;

4) the reference period for the data.

b) the collection timeline, including milestones for data acquisition, validation, and integration;

c) the collection methodology, specifying:

1) whether data is obtained through direct collection, automated retrieval, or third-party provision;

2) any data sharing agreements or access arrangements required;

3) the protocols for handling data access denials or partial responses.

d) the roles and responsibilities of personnel involved in data collection.

7.2.2 Data Collection Execution

Data collection shall be executed in accordance with the collection plan. The ranking entity shall:

a) collect data for all ranked entities within the defined reference period;

b) apply consistent collection procedures across all entities;

c) record the date and method of acquisition for each data point;

d) flag any deviations from the collection plan and document the reasons and implications.

7.2.3 Data Quality Verification

Upon completion of data collection, the ranking entity shall conduct quality verification as specified in 7.3. Data that does not meet the quality standards shall be:

a) returned to the source for correction, where feasible;

b) supplemented with data from alternative sources;

c) marked as missing and handled in accordance with 7.4.1.

7.2.4 Data Entry and Labelling

Data that has passed quality verification shall be entered into the ranking database with the following annotations:

a) the entity identifier;

b) the indicator identifier;

c) the data source identifier and tier level;

d) the reference period;

e) the acquisition date;

f) any quality flags or caveats;

g) the version identifier of the data collection cycle.


7.3 Data Quality Standards

Data used in ranking systems shall meet the quality standards defined in 7.3.1 through 7.3.5. These standards are aligned with the principles of ISO 8000 (Data Quality).

7.3.1 Accuracy

Data shall accurately represent the real-world value it purports to measure.

The accuracy of data shall be assessed by:

a) comparing data values against independent, authoritative sources where available;

b) performing plausibility checks (e.g., verifying that values fall within expected ranges);

c) verifying computational derivations (e.g., that ratios are correctly computed from their numerator and denominator).

The ranking entity shall document the accuracy assessment methods and results for each indicator.

7.3.2 Completeness

The data set shall contain values for all entities and all indicators within the ranking scope, or shall explicitly document the extent and pattern of missing data.

Completeness shall be measured as:

\[\text{Completeness} = \frac{\text{Number of non-missing data points}}{\text{Total number of expected data points}} \times 100\%\]

The ranking entity shall:

a) report the overall completeness rate for each indicator;

b) report the completeness rate for each entity;

c) document the pattern of missing data (e.g., random, systematic, or concentrated in specific entity groups);

d) assess whether the pattern of missing data introduces bias (see 4.7).

7.3.3 Consistency

Data shall be internally consistent and consistent across sources.

Consistency shall be verified by:

a) cross-referencing data values from different sources for the same entity and indicator;

b) checking that aggregate values (e.g., totals, averages) are consistent with their component values;

c) verifying that units of measurement, currencies, and temporal references are uniform across the dataset.

Where inconsistencies are identified, the ranking entity shall:

1) investigate the cause of the inconsistency;

2) determine the correct value through reference to the most authoritative source;

3) document the inconsistency and the resolution.

7.3.4 Timeliness

Data shall be sufficiently current for the ranking purpose.

The ranking entity shall:

a) specify the maximum acceptable age of data for each indicator, based on the rate of change of the underlying phenomenon and the purpose of the ranking;

b) report the reference date of each data point;

c) document any data that exceeds the maximum acceptable age and the justification for its continued use.

7.3.5 Traceability

Each data point shall be traceable to its original source through a documented chain of custody.

Traceability shall be maintained by:

a) recording the source, acquisition method, and date for each data point;

b) preserving the original data in its unmodified form;

c) maintaining a log of all transformations applied to the data (see 7.4).


7.4 Data Processing Specifications

7.4.1 Missing Data Handling

The ranking entity shall specify and document the method for handling missing data. The following methods are recognised:

a) Imputation: Missing values may be replaced with estimated values. The imputation method shall be documented and shall include:

1) Mean/median imputation: Replacing missing values with the mean or median of the available values for that indicator across other entities. This method shall only be used when the proportion of missing data is less than 5 % and the data is missing at random.

2) Regression imputation: Estimating missing values from a regression model using correlated indicators as predictors. The regression model and its goodness-of-fit shall be documented.

3) Multiple imputation: Generating multiple plausible values for each missing data point and combining the results. This method should be used when the proportion of missing data exceeds 5 %.

b) Exclusion: Entities with missing data for a specific indicator may be excluded from the calculation of that indicator. The ranking entity shall:

1) document the number and identity of excluded entities;

2) assess whether exclusion introduces systematic bias;

3) report the impact of exclusion on the ranking results.

c) Flagging: Missing values may be flagged without imputation or exclusion, and the indicator score may be calculated on the basis of available data. The ranking entity shall:

1) clearly label the flag in the data and results;

2) disclose the proportion of flagged data for each entity;

3) report the impact of flagging on the ranking results.

7.4.2 Outlier Handling

The ranking entity shall establish and document procedures for identifying and handling outliers.

7.4.2.1 Outlier Identification

Outliers shall be identified using at least one of the following methods:

a) Statistical methods:

1) Z-score method: data points with z > 3 shall be flagged as potential outliers.

2) Interquartile range (IQR) method: data points below Q1 − 1.5 × IQR or above Q3 + 1.5 × IQR shall be flagged as potential outliers.

3) Mahalanobis distance: for multivariate outlier detection, observations with Mahalanobis distance exceeding the χ² critical value at α = 0.01 shall be flagged.

b) Domain knowledge: Domain experts may identify outliers based on substantive knowledge of the data context.

7.4.2.2 Outlier Treatment

Identified outliers shall be handled by one of the following methods:

a) Retention with justification: If the outlier is determined to be a genuine data point reflecting a real phenomenon, it shall be retained with documented justification.

b) Winsorisation: Outlying values shall be set to the value of a specified percentile (e.g., the 1st and 99th percentiles). The percentile thresholds shall be documented.

c) Transformation: The data may be transformed (e.g., log transformation) to reduce the influence of outliers. The transformation shall be documented and its impact on the distribution shall be assessed.

d) Exclusion: Outliers may be excluded only when they are determined to be data errors. The exclusion shall be documented with the reason.

7.4.2.3 Outlier Documentation

All outlier identification and treatment decisions shall be documented in the data processing log, including:

a) the identification method and criteria used;

b) the data points identified as outliers;

c) the treatment applied and the justification;

d) the impact on the ranking results.

7.4.3 Data Standardisation and Normalisation

Where indicators are measured on different scales, the data shall be standardised or normalised before aggregation. The methods specified in Chapter 8, Section 8.2 shall apply.

The ranking entity shall:

a) specify the standardisation or normalisation method for each indicator;

b) document the reason for the choice of method;

c) verify that the chosen method does not introduce bias or distort the indicator distribution in a manner that disadvantages specific entity groups.


7.5 Data Ethics and Privacy

7.5.1 Personal Information Protection

Where the ranking system collects or processes personal data, the ranking entity shall comply with applicable data protection legislation, including but not limited to the General Data Protection Regulation (GDPR) where it applies.

The ranking entity shall:

a) establish a lawful basis for the collection and processing of personal data;

b) minimise the collection of personal data to what is necessary for the ranking purpose;

c) implement appropriate technical and organisational measures to protect personal data;

d) ensure that personal data is not used for purposes incompatible with the ranking purpose.

7.5.2 Data Use Authorisation

The ranking entity shall ensure that it has the right to use all data incorporated in the ranking system. The ranking entity shall:

a) verify the terms of use or licence for each data source;

b) obtain necessary permissions for data that is not publicly available;

c) comply with any restrictions on the use, reproduction, or redistribution of data.

7.5.3 Data Awareness and Objection Rights of Ranked Entities

Ranked entities shall have the following rights with respect to data about them that is used in the ranking:

a) Right to be informed: Ranked entities shall be informed of the data used about them in the ranking, upon request.

b) Right to object: Ranked entities shall have the right to challenge the accuracy of data used about them. The ranking entity shall:

1) establish a documented procedure for receiving and processing data objections;

2) investigate the objection within a reasonable period (not exceeding 30 calendar days);

3) correct any data found to be inaccurate and document the correction;

4) inform the objecting entity of the outcome of the investigation.

c) Right to correct: Where data is found to be inaccurate, the ranking entity shall correct the data and recompute the affected ranking results.

Note: The right to object does not extend to the methodological choices of the ranking system (e.g., indicator selection, weighting), which are governed by the principles in Chapter 4 and the procedures in Chapters 5 and 6.


7.6 Data Audit and Archival

7.6.1 Data Provenance Chain

The ranking entity shall maintain a complete provenance chain for all data used in the ranking. The provenance chain shall include:

a) the original source of each data point;

b) the date and method of acquisition;

c) all transformations, corrections, and adjustments applied to the data;

d) the personnel responsible for each transformation;

e) the date and authorisation of each transformation.

7.6.2 Raw Data Archival Requirements

The ranking entity shall archive all raw (unprocessed) data used in the ranking for a minimum period of five years from the date of publication of the ranking results. The archived data shall:

a) be stored in a secure, accessible format;

b) include all metadata required for traceability (see 7.3.5);

c) be protected against unauthorised modification;

d) be available for audit purposes upon request.

7.6.3 Data Change Log

The ranking entity shall maintain a data change log that records all modifications to the ranking dataset after initial collection. Each log entry shall include:

a) the date and time of the change;

b) the identity of the person authorising the change;

c) a description of the change, including the original and modified values;

d) the reason for the change;

e) the impact of the change on the ranking results.

The data change log shall be retained for the same period as the raw data (see 7.6.2).


Note: The data quality standards in this chapter are aligned with ISO 8000 (Data Quality) and ISO 9001 (Quality Management Systems). Ranking entities that maintain certification under these standards may reference their certified processes as evidence of compliance with the relevant provisions of this chapter.