Team Ai
Modelpublic

CodeHima/TOSRobertaV2

sourceHugging Facemitupdated 2y agoView on Hugging Face
0likes12downloads
README.md124 linesDownload Raw Back to root
1---2license: mit3language:4- en5widget:6  - text: "You have the right to use CommunityConnect for its intended purpose of connecting with others, sharing content responsibly, and engaging in constructive dialogue. You are responsible for the content you post and must respect the rights and privacy of others."7    example_title: "Fair Clause"8  - text: " We reserve the right to suspend, terminate, or restrict your access to the platform at any time and for any reason, without prior notice or explanation. This includes but is not limited to violations of our community guidelines or terms of service, as determined solely by ConnectWorld."9    example_title: "Unfair Clause"10metrics:11- accuracy12- precision13- f114- recall15library_name: transformers16pipeline_tag: text-classification17---18# TOSRobertaV2: Terms of Service Fairness Classifier19 20## Model Description21 22TOSRobertaV2 is a fine-tuned RoBERTa-large model designed to classify clauses in Terms of Service (ToS) documents based on their fairness level. The model categorizes clauses into three classes: clearly fair, potentially unfair, and clearly unfair.23 24## Intended Use25 26This model is intended for:27- Analyzing Terms of Service documents for potential unfair clauses28- Assisting legal professionals in reviewing contracts29- Helping consumers understand the fairness of agreements they're entering into30- Supporting researchers studying fairness in legal documents31 32## Training Data33 34The model was trained on the CodeHima/TOS_DatasetV3, which contains labeled clauses from various Terms of Service documents.35 36## Training Procedure37 38- Base model: RoBERTa-large39- Training type: Fine-tuning40- Number of epochs: 541- Optimizer: AdamW42- Learning rate: 2e-543- Batch size: 844- Weight decay: 0.0145- Training loss: 0.385197297365252946 47## Evaluation Results48 49### Validation Set Performance50 51- Accuracy: 0.8652- F1 Score: 0.858853- Precision: 0.859854- Recall: 0.860055 56### Test Set Performance57 58- Accuracy: 0.865159 60### Training Progress61 62| Epoch | Training Loss | Validation Loss | Accuracy | F1     | Precision | Recall  |63|-------|---------------|-----------------|----------|--------|-----------|---------|64| 1     | 0.5391        | 0.493973        | 0.798095 | 0.7997 | 0.8056    | 0.79810 |65| 2     | 0.4621        | 0.489970        | 0.831429 | 0.8320 | 0.8330    | 0.83143 |66| 3     | 0.3954        | 0.674849        | 0.821905 | 0.8250 | 0.8349    | 0.82191 |67| 4     | 0.3783        | 0.717495        | 0.860000 | 0.8588 | 0.8598    | 0.86000 |68| 5     | 0.1542        | 0.881050        | 0.847619 | 0.8490 | 0.8514    | 0.84762 |69 70## Limitations71 72- The model's performance may vary on ToS documents from domains or industries not well-represented in the training data.73- It may struggle with highly complex or ambiguous clauses.74- The model's understanding of "fairness" is based on the training data and may not capture all nuances of legal fairness.75 76## Ethical Considerations77 78- This model should not be used as a substitute for professional legal advice.79- There may be biases present in the training data that could influence the model's judgments.80- Users should be aware that the concept of "fairness" in legal documents can be subjective and context-dependent.81 82## How to Use83 84You can use this model directly with the Hugging Face `transformers` library:85 86```python87from transformers import AutoTokenizer, AutoModelForSequenceClassification88import torch89 90tokenizer = AutoTokenizer.from_pretrained("CodeHima/TOSRobertaV2")91model = AutoModelForSequenceClassification.from_pretrained("CodeHima/TOSRobertaV2")92 93text = "Your clause here"94inputs = tokenizer(text, return_tensors="pt", truncation=True, padding=True, max_length=512)95 96with torch.no_grad():97    logits = model(**inputs).logits98 99probabilities = torch.softmax(logits, dim=1)100predicted_class = torch.argmax(probabilities, dim=1).item()101 102classes = ['clearly fair', 'potentially unfair', 'clearly unfair']103print(f"Predicted class: {classes[predicted_class]}")104print(f"Probabilities: {probabilities[0].tolist()}")105```106 107## Citation108 109If you use this model in your research, please cite:110 111```112@misc{TOSRobertaV2,113  author = {CodeHima},114  title = {TOSRobertaV2: Terms of Service Fairness Classifier},115  year = {2024},116  publisher = {Hugging Face},117  journal = {Hugging Face Model Hub},118  howpublished = {\url{https://huggingface.co/CodeHima/TOSRobertaV2}}119}120```121 122## License123 124This model is released under the MIT license.