Team Ai
Modelpublic

gplsi/Aitana-ClearLangDetection-R-1.0

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
0likes11downloads
Model Card

Aitana-ClearLangDetection-R-1.0

Table of Contents

Model Description

This model is fine-tuned from BSC-LT/mRoBERTa for the task of clear language classification in Spanish texts.

It predicts among three categories of linguistic clarity:

  • —TXT: Original text
  • —FAC: Facilitated text
  • —LF: Easy-to-read text

Training Details

Training Data

The dataset consists of Spanish texts annotated with clarity levels:

  • —Training set: 9,299 instances
  • —Test set: 3,723 instances
  • —Extra test set: 465 instances (texts from non-contiguous categories not seen during training, used to evaluate generalization)

The original dataset is available in gplsi/discriminative_clearsim_es.

Training Hyperparameters

  • —learning_rate: 2e-5
  • —numtrainepochs: 2
  • —perdevicetrainbatchsize: 8
  • —perdeviceevalbatchsize: 8
  • —overwriteoutputdir: true
  • —logging_strategy: steps
  • —logging_steps: 10
  • —seed: 852
  • —fp16: true

Evaluation

Combined test set (4,188 instances)

Confusion Matrix

Pred FACPred LFPred TXT
True FAC1373158
True LF2913670
True TXT1611379
ClassPrecisionRecallF1-scoreSupport
FAC0.96830.98350.97581396
LF0.98840.97920.98381396
TXT0.99420.98780.99101396
  • —Accuracy: 0.9835
  • —Macro Avg F1: 0.9836 ---

Test set (3,723 instances)

Confusion Matrix

Pred FACPred LFPred TXT
True FAC1220138
True LF2812130
True TXT1311227
ClassPrecisionRecallF1-scoreSupport
FAC0.96750.98310.97521241
LF0.98860.97740.98301241
TXT0.99350.98870.99111241
  • —Accuracy: 0.9831
  • —Macro Avg F1: 0.9831 ---

Extra test set (465 instances)

Confusion Matrix

Pred FACPred LFPred TXT
True FAC15320
True LF11540
True TXT30152
ClassPrecisionRecallF1-scoreSupport
FAC0.97450.98710.9808155
LF0.98720.99360.9903155
TXT1.00000.98060.9902155
  • —Accuracy: 0.9871
  • —Macro Avg F1: 0.9871

Technical Specifications

Hardware and Software

For training, we used custom code developed to fine-tuned model using Transformers library.

Compute Infrastructure

This model was trained on NVIDIA DGX systems equipped with A100 GPUs, which enabled efficient large-scale training. For this model, we used one A100 GPU.

Additional Information

Author

The model has been developed by the Language and [Information Systems Group (GPLSI)](https://gplsi.dlsi.ua.es/) and the [Centro de Inteligencia Digital (CENID)](https://cenid.es), both part of the [University of Alicante (UA)](https://www.ua.es/es/), as part of their ongoing research in Natural Language Processing (NLP).

Funding

This work is funded by the Ministerio para la Transformación Digital y de la Función Pública, co-financed by the EU – NextGenerationEU, within the framework of the project Desarrollo de Modelos ALIA.

Acknowledgments

We would like to express our gratitude to all individuals and institutions that have contributed to the development of this work.

Special thanks to:

We also acknowledge the financial, technical, and scientific support of the Ministerio para la Transformación Digital y de la Función Pública - Funded by EU – NextGenerationEU within the framework of the project Desarrollo de Modelos ALIA, whose contribution has been essential to the completion of this research.

License

Apache License, Version 2.0

Disclaimer

This model has been developed and fine-tuned specifically for classification task. The authors are not responsible for potential errors, misinterpretations, or inappropriate use of the model beyond its intended purpose.

Reference

If you use this model in your research or work, please cite it as follows:

bibtex
@misc{gplsi-aitama-clear-r-1.0,
  author       = {Sepúlveda-Torres, Robiert and Martínez-Murillo, Iván and Bonora, Mar and Consuegra-Ayala, Juan Pablo and Galeano, Santiago and Miró Maestre, María and  and Grande, Eduardo and Canal-Esteve, Miquel and Estevanell-Valladares, Ernesto L. and Yáñez-Romero, Fabio and Gutierrez, Yoan and Abreu Salas, José Ignacio and Lloret, Elena and Montoyo, Andrés and Muñoz-Guillena and Palomar, Manuel},
  title        = {Aitana-ClearLangDetection-R-1.0: Fine-tuned model for clear language classification (TXT, FAC, LF)},
  year         = {2025},
  institution  = {Language and Information Systems Group (GPLSI) and Centro de Inteligencia Digital (CENID), University of Alicante (UA)},
  howpublished = {\url{https://huggingface.co/gplsi/Aitana-ClearLangDetection-R-1.0}},
  note         = {Accessed: 2025-10-03}
}

Copyright © 2026 Language and Information Systems Group (GPLSI) and Centro de Inteligencia Digital (CENID), University of Alicante (UA). Distributed under the Apache License 2.0.