agentlans/mdeberta-v3-base-readability
DeBERTa V3 Base for Multilingual Readability Assessment
This is a fine-tuned version of the multilingual DeBERTa model (mdeberta) for assessing text readability across languages.
Model Details
- Architecture: mdeberta-base
- Task: Regression (Readability Assessment)
- Training Data: agentlans/tatoeba-english-translations dataset containing 39 100 English translations
- Input: Text in any of the supported languages by DeBERTa
- Output: Estimated U.S. grade level for text comprehension
- higher values indicate more complex text
Performance
Root mean squared error (RMSE) on 20% held-out validation set: 1.063
Training Data
The model was trained on agentlans/tatoeba-english-translations.
Usage
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model_name="agentlans/mdeberta-v3-base-readability"
# Put model on GPU or else CPU
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = model.to(device)
def readability(text):
"""Processes the text using the model and returns its logits.
In this case, it's reading grade level in years of education
(the higher the number, the harder it is to read the text)."""
inputs = tokenizer(text, return_tensors="pt", truncation=True, padding=True).to(device)
with torch.no_grad():
logits = model(**inputs).logits.squeeze().cpu()
return logits.tolist()
readability("Your text here.")Results
In this study, 10 English text samples of varying readability were generated and translated into Arabic, Chinese, French, Russian, and Spanish using Google Translate. This resulted in a total of 50 translated samples, which were subsequently analyzed by a trained classifier to predict their readability scores.
<details> <summary>The following table presents the 10 original texts along with their translations:</summary>
</details>
The scatterplot below illustrates the predicted readability scores grouped by each text sample. Notably, the prediction scores exhibit low variability across different languages for the same text, indicating a consistent assessment of translation readability regardless of the target language.
<img src="plot.png" alt="Scatterplot of predicted quality scores grouped by text sample and language" width="100%"/>
This analysis highlights the effectiveness of using machine learning classifiers in evaluating textual readability across multiple languages.
Limitations
- Performance may vary for texts significantly different from the training data
- Output is based on statistical patterns and may not always align with human judgment
- Readability is assessed purely on textual features, not considering factors like subject familiarity or cultural context
Ethical Considerations
- Should not be used as the sole determinant of text suitability for specific audiences
- Results may reflect biases present in the training data sources
- Care should be taken when using these models in educational or publishing contexts
