mljn/unga-climate-classifier
unga-climate-classifier
A fine-tuned version of microsoft/deberta-v3-base that classifies whether an English sentence from a United Nations General Assembly (UNGA) speech is about climate change.
It was developed to build the Executive Climate Change Attention (ECCA) indicator, introduced in Grossman and Jan (2026), Global Environmental Politics. The indicator measures the share of climate-related sentences in the annual speeches that governments of 195 countries have delivered at the UNGA since 1985.
Labels
The model returns one of two labels for each sentence:
climate: the sentence is about climate changeother: it is not
Definition
A sentence is coded as climate-related when it explicitly refers to anthropogenic climate change, including its causes and consequences, related international and national events, and policies addressing it. Broader references to environmental degradation or energy policy (e.g. renewable energy) are not coded as climate-related unless they mention climate change. Extreme weather events are coded only when attributed to climate change.
Training data
5,600 manually annotated sentences from the UN General Debate Corpus (Baturo, Dasandi and Mikhaylov 2017), which contains the English versions of the speeches. Intercoder reliability: Krippendorff's α = 0.795.
After removing duplicates, the 5,621 sentences were split into training (3,935), validation (843) and test (843) sets, stratified by label. About 31% of sentences in each set are labelled climate.
How to use
from transformers import pipeline
classifier = pipeline("text-classification", model="mljn/unga-climate-classifier")
classifier("Climate change poses a fundamental threat to our future.")
# [{'label': 'climate', 'score': 0.9988}]From sentences to the ECCA indicator
The ECCA indicator is the share of sentences in a speech that the model classifies as climate-related. To reproduce it, split each speech into sentences, classify every sentence, and compute the proportion labelled climate.
Limitations
- The model was trained and evaluated on English sentences from UNGA speeches. Its performance on other kinds of text, such as manifestos, parliamentary debates or social media posts, has not been tested.
- It classifies single sentences. Longer texts should be split into sentences first; inputs are truncated at 512 tokens.
- It identifies explicit references to climate change only, following the definition above. Implicit references, such as energy or environmental policy that does not mention climate change, are classified as
otherby design.
Evaluation
Results on the held-out evaluation set:
The model outperforms a random forest classifier and several keyword dictionaries (see the paper, Figure 1).
<details> <summary>Training details</summary>
Training hyperparameters
- learning_rate: 2e-05
- trainbatchsize: 16
- evalbatchsize: 80
- seed: 42
- gradientaccumulationsteps: 2
- totaltrainbatch_size: 32
- optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
- lrschedulertype: linear
- lrschedulerwarmup_ratio: 0.06
- num_epochs: 5
Training results
Framework versions
- Transformers 4.36.2
- Pytorch 2.5.0+cu121
- Datasets 2.6.0
- Tokenizers 0.15.2
</details>
Data and replication
The ECCA indicator and replication materials are available at github.com/malojan/executive_climate_change_attention.
Citation
If you use this model, the underlying dataset or the indicator, please cite:
Grossman, Emiliano, and Malo Jan. 2026. "Executive Climate Change Attention: Toward an Indicator of Comparative Climate Change Attention." Global Environmental Politics 26 (2): 137–150. https://doi.org/10.1162/glep.a.1
@article{grossman2026executive,
title = {Executive Climate Change Attention: Toward an Indicator of Comparative Climate Change Attention},
author = {Grossman, Emiliano and Jan, Malo},
journal = {Global Environmental Politics},
volume = {26},
number = {2},
pages = {137--150},
year = {2026},
doi = {10.1162/glep.a.1}
}