Team Ai
Modelpublic

MatteoFasulo/xlm-roberta-xstance

sourceHugging Facemitupdated 3mo agoView on Hugging Face
1likes54downloads
Model Card

<a href="https://huggingface.co/spaces/MatteoFasulo/huggingface-static-200e93" target="_blank"> <img src="https://raw.githubusercontent.com/gradio-app/trackio/refs/heads/main/trackio/assets/badge.png" alt="Visualize in Trackio" title="Visualize in Trackio" style="height:40px;"> </a>

xlm-roberta-xstance

A multilingual stance detection model fine-tuned from FacebookAI/xlm-roberta-base on the ZurichNLP/x_stance dataset.

The model predicts whether a political comment expresses a FAVOR or AGAINST stance toward a given political question. It supports multilingual inference and demonstrates strong cross-lingual transfer across Swiss national languages.


Highlights

  • —🌍 Multilingual stance detection (🇩🇪 German (75%), 🇫🇷 French (25%), and 🇮🇹 Italian (only a few to test zero-shot cross-lingual transfer))
  • —⚡ Built on XLM-RoBERTa
  • —🎯 Binary stance classification (FAVOR / AGAINST)
  • —🔄 Cross-lingual transfer capabilities

Performance

Validation Split

Evaluation on the validation split:

MetricScore
Loss0.5225
Accuracy76.87%
Macro F176.87%

X-Stance Test Set

Detailed evaluation on the X-Stance test set as provided in the original repository:

Evaluation settingLanguageMacro F1
New commentsGerman (DE)77.348
New commentsFrench (FR)76.914
New questionsGerman (DE)71.334
New questionsFrench (FR)70.935
New topicsGerman (DE)72.866
New topicsFrench (FR)75.659
New commentsItalian (IT)73.768

The evaluation script used to obtain these results is:

bash
python evaluate.py \
  --gold data/test.jsonl \
  --pred predictions/xlm_roberta_pred.jsonl
Note: xlm_roberta_pred.jsonl is provided in this repository for reproducibility. The evaluation script is available in the original repository at https://github.com/ZurichNLP/xstance .

Quick Start

Installation

bash
pip install transformers torch

Run inference

Using the pipeline API (Recommended)

python
from transformers import pipeline

classifier = pipeline(
    task="text-classification",
    model="MatteoFasulo/xlm-roberta-xstance"
)

question = "Soll der Bundesrat ein Freihandelsabkommen mit den USA anstreben?"

comment = "Nicht unter einem Präsidenten, welcher die Rechte anderer mit Füssen tritt und Respektlos gegenüber ändern ist."

result = classifier(
    {
        "text": question,
        "text_pair": comment,
    }
)

print(result)

Example output:

python
[{'label': 'AGAINST', 'score': 0.9823}]

For sequence-pair classification tasks such as stance detection, the text-classification pipeline accepts a dictionary with "text" and "text_pair" keys.


Using AutoModelForSequenceClassification

python
import torch
from transformers import (
    AutoTokenizer,
    AutoModelForSequenceClassification,
)

model_name = "MatteoFasulo/xlm-roberta-xstance"

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)

question = "Soll der Bundesrat ein Freihandelsabkommen mit den USA anstreben?"

comment = "Nicht unter einem Präsidenten, welcher die Rechte anderer mit Füssen tritt und Respektlos gegenüber ändern ist."

inputs = tokenizer(
    question,
    comment,
    return_tensors="pt",
    truncation=True,
)

with torch.no_grad():
    outputs = model(**inputs)

probabilities = torch.softmax(outputs.logits, dim=-1)

prediction = probabilities.argmax(dim=-1).item()

id2label = model.config.id2label

print("Prediction:", id2label[prediction])
print("Confidence:", probabilities[0, prediction].item())

Example output:

text
Prediction: AGAINST
Confidence: 0.9823

Input Format

The model expects two text sequences:

  1. 1.Target political question
  2. 2.Candidate comment

Example:

Question:
Should Switzerland increase renewable energy subsidies?

Comment:
Investing in renewable energy will reduce emissions and improve energy independence.

The tokenizer automatically formats these as sentence pairs for XLM-RoBERTa.


Output Labels

The classifier predicts one of two classes.

LabelDescription
FAVORThe comment supports the target question.
AGAINSTThe comment opposes the target question.

The model outputs logits for both classes.


Model Description

This model is a fine-tuned version of FacebookAI/xlm-roberta-base trained for multilingual, multi-target stance detection.

Unlike sentiment analysis, stance detection predicts whether a text supports or opposes a specific target question.

Because the underlying encoder is multilingual, the model can transfer knowledge across languages and perform inference on languages that were only partially represented during training.


Intended Uses

The model is suitable for:

  • —Political stance detection
  • —Cross-lingual stance classification
  • —Research on multilingual NLP
  • —Opinion mining
  • —Benchmarking stance detection methods

Out-of-Scope Uses

This model is not intended for:

  • —Fact checking
  • —Political affiliation prediction
  • —Hate speech detection
  • —Toxicity classification
  • —General sentiment analysis
  • —Automated political decision-making

Training Dataset

Training was performed using the ZurichNLP/x_stance dataset.

Dataset characteristics:

  • —150+ political questions
  • —67,000 candidate comments
  • —Swiss political debates
  • —Multilingual annotations

Languages:

  • —German (majority)
  • —French
  • —Italian

Each sample consists of:

(Target Question, Candidate Comment)
→ FAVOR / AGAINST

Training Procedure

Hyperparameters

ParameterValue
Base modelFacebookAI/xlm-roberta-base
Learning rate2e-5
Batch size16
Epochs3
OptimizerAdamW (Torch Fused)
SchedulerLinear
Warmup steps850
Mixed precisionNative AMP
Seed42

Training Results

Training LossEpochStepValidation LossAccuracyMacro F1
0.5537128530.57490.71750.7175
0.4712257060.49570.75880.7587
0.3804385590.52250.76870.7687

Limitations

Although the model performs well on multilingual political stance detection, several limitations should be considered.

  • —Trained primarily on Swiss political debates.
  • —Binary labels only (no Neutral class).
  • —Performance outside politics has not been evaluated.
  • —Implicit or sarcastic opinions remain challenging.
  • —Domain shift may reduce performance on social media or informal discussions.

Ethical Considerations

This model predicts stance, not factual correctness.

Predictions should not be interpreted as:

  • —political affiliation
  • —truthfulness
  • —misinformation detection
  • —ideological profiling

Human oversight is recommended for any downstream application.


Framework Versions

  • —Transformers 5.12.1
  • —PyTorch 2.8.0
  • —Datasets 5.0.0
  • —Tokenizers 0.22.2

Citation

If you use this model, please cite the original X-Stance dataset.

bibtex
@inproceedings{vamvas2020xstance,
    author    = "Vamvas, Jannis and Sennrich, Rico",
    title     = "{X-Stance}: A Multilingual Multi-Target Dataset for Stance Detection",
    booktitle = "Proceedings of the 5th Swiss Text Analytics Conference (SwissText) \& 16th Conference on Natural Language Processing (KONVENS)",
    address   = "Zurich, Switzerland",
    year      = "2020",
    month     = "jun",
    url       = "http://ceur-ws.org/Vol-2624/paper9.pdf"
}

License

This model is released under the MIT License.

Please also respect the licenses of:

  • —FacebookAI/xlm-roberta-base
  • —ZurichNLP/x_stance

Acknowledgements

This model builds upon:

  • —Facebook AI Research for XLM-RoBERTa
  • —Zurich NLP Group for the X-Stance dataset
  • —Hugging Face Transformers