Team Ai
Datasetpublic

lszoszk/uhri-recommendations

UN Human Rights Recommendations (UHRI+) A cleaned and normalized dataset of 272,502 country-specific human rights observations and recommendations from three UN mechanisms, covering 2006–2026. Built on the OHCHR Universal Human Rights Index (UHRI) — the authoritative UN source — with systematic data quality improvements applied on top. Coverage Mechanism Records Bodies Universal Periodic Review (UPR) 130,870 1 Treaty Bodies 120,990 12 (CCPR, CAT, CERD… See the full description on the dataset page: https://huggingface.co/datasets/lszoszk/uhri-recommendations.

sourceHugging Facecc-by-nc-4.0updated 9d agoView on Hugging Face
0likes147downloads
Dataset Card

UN Human Rights Recommendations (UHRI+)

A cleaned and normalized dataset of 272,502 country-specific human rights observations and recommendations from three UN mechanisms, covering 2006–2026.

Built on the OHCHR Universal Human Rights Index (UHRI) — the authoritative UN source — with systematic data quality improvements applied on top.

Coverage

MechanismRecordsBodies
Universal Periodic Review (UPR)130,8701
Treaty Bodies120,99012 (CCPR, CAT, CERD, CEDAW, CRC, CESCR, CMW, CRPD, CED, SPT, CRC-OP-AC, CRC-OP-SC)
Special Procedures (SR, IE, WG)20,64257

Totals are exact counts over this dataset's body field (130,870 + 120,990 + 20,642 = 272,502) and can be reproduced with collections.Counter(ds["train"]["body"]).

199 countries and entities covered, including Cook Islands, Niue (non-UN-member states), State of Palestine\, Holy See (UN observer states), Kosovo\ (contested status) and the European Union (a regional bloc, not a State).

Cleaning pipeline (5 stages)

The canonical pipeline description (with full detail and examples) lives on the dashboard's Methodology page. Summary:

StageWhat it doesScale
1Rule-based OCR & HTML repair (no AI): stray tags, mid-word OCR splits (develop- ment → development), whitespace, broken citationsStages 1–4 together edited ~56,000 records (≈21 % of the dataset)
2LLM-assisted residue review for hard OCR corruption (bounded prompt, structural validation)412 records (0.15 %)
3Deterministic annotation_type normalisation (multilingual regex classifier, audit-tagged)2,603 of ~3,294 missing/UUID-coded types re-labelled
4Country backfill from the UN document symbol (CEDAW/C/ALB/CO/4 → Albania)106 records
5Artefact drop — rows with no citable content (empty placeholders, duplicate annotation IDs, HTML-only scaffolding)11 rows (272,513 → 272,502)

text_cleaned reflects Stages 1–2 (text repair); Stages 3–5 act on metadata and row selection. text_original preserves the raw OHCHR export text.

Applied data fixes (included in this parquet): body-name normalisation merged duplicate issuing-body spellings (six treaty bodies and several Special-Procedure mandates appeared under both short codes and long names upstream) and resolved 20 blank body values from their document symbols — 1,481 rows changed, 78 → 70 distinct bodies. The same fixes are folded into the live dashboard's data pipeline.

September 2026 refresh (v2026.09): adds 4,831 records — documents OHCHR published to the Index after the previous snapshot, including a full Universal Periodic Review session — and corrects the existing ones: 891 rows that carried a bare ISO-2 code as the country (e.g. SG) now carry the full name (Singapore), and 1,587 `citation` strings that still showed pre-normalisation body names or lacked a backfilled country were regenerated from the corrected fields. All other columns are unchanged for the 267,671 rows of the previous snapshot.

Schema

FieldTypeDescription
annotation_idstringOHCHR UUID
symbolstringUN document symbol (e.g. CEDAW/C/ALB/CO/4)
bodystringIssuing body short code (e.g. CEDAW, UPR, SR Torture)
publication_datestringISO date (YYYY-MM-DD)
publication_yearintYear for fast filtering
annotation_typestringRecommendations or Concerns/Observations
countrieslist[str]Target countries (full names)
regionslist[str]UN M49 regions
affected_personslist[str]OHCHR affected-person categories
themeslist[str]OHCHR thematic categories
sdgslist[str]SDG tags
text_originalstringRaw OHCHR export text
text_cleanedstringText after 5-stage cleaning pipeline
citationstringFormatted citation

Quick start

python
from datasets import load_dataset

ds = load_dataset("lszoszk/uhri-recommendations")

# Filter by country
poland = ds["train"].filter(lambda x: "Poland" in x["countries"])

# Filter by mechanism
upr = ds["train"].filter(lambda x: x["body"] == "UPR")

# Search recommendations
torture = ds["train"].filter(lambda x: "torture" in x["text_cleaned"].lower())

Live dashboard

An interactive search and analytics dashboard built on this dataset: [https://lszoszk.github.io/UnitedNations_recommendations/](https://lszoszk.github.io/UnitedNations_recommendations/)

Features: boolean + wildcard + FTS5 stemmed search, per-country profiles, hex map, timeline analytics, SDG mapping.

Citation

bibtex
@dataset{szoszkiewicz2026uhri,
  author    = {Szoszkiewicz, Lukasz},
  title     = {UN Human Rights Recommendations (UHRI+)},
  year      = {2026},
  version   = {v2026.09},
  doi       = {10.5281/zenodo.21319464},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/datasets/lszoszk/uhri-recommendations},
  note      = {Cleaned dataset built on OHCHR Universal Human Rights Index, 2006--2026. The DOI is the concept DOI of the UHRI+ project on Zenodo.}
}

Licence

Source data © OHCHR — shared freely and without restriction per the [UHRI data policy](https://uhri.ohchr.org/en/our-data-api).

⚠️ The OHCHR data policy may change. Always consult the current terms at uhri.ohchr.org/en/our-data-api before redistributing. This dataset card reflects the policy as of 2026-05-18.

Cleaning pipeline, normalisation, Stage 4 country backfill, and all metadata additions © L. Szoszkiewicz (Adam Mickiewicz University, Poznań), released under CC BY-NC 4.0 — free for research and education, commercial use requires permission.

Card changelog

  • —2026-10-01 — refreshed to v2026.09: 272,502 records (+4,831), 891 country codes and 1,587 citations corrected (see Applied data fixes), shards renamed train-0000N-of-00005. Licence unchanged.
  • —2026-06-11 — corrected the mechanism totals (Treaty Bodies 119,532; Special Procedures 20,514 — previously listed as rough estimates that were each ~16k off) and replaced the pipeline summary with the canonical five-stage description matching the published methodology.
  • —2026-05-18 — initial release.