Omartificial-Intelligence-Space/Arabic-NLi-Pair
Arabic-NLI-PAir Dataset Summary The Arabic Version of SNLI and MultiNLI datasets. (Pair Subset) Originally used for Natural Language Inference (NLI), Dataset may be used for training/finetuning an embedding model for semantic textual similarity. Pair Subset Columns: "anchor", "positive" Column types: str, str Examples: { "anchor": "كيف أكون جيولوجياً جيداً؟", "positive": "ماذا علي أن أفعل لأكون جيولوجياً عظيماً؟" } Disclaimer… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-NLi-Pair.
Arabic-NLI-PAir
Dataset Summary
- The Arabic Version of SNLI and MultiNLI datasets. (Pair Subset)
- Originally used for Natural Language Inference (NLI),
- Dataset may be used for training/finetuning an embedding model for semantic textual similarity.
Pair Subset
- Columns: "anchor", "positive"
- Column types: str, str
Examples:
{
"anchor": "كيف أكون جيولوجياً جيداً؟",
"positive": "ماذا علي أن أفعل لأكون جيولوجياً عظيماً؟"
}Disclaimer
Please note that the translated sentences are generated using neural machine translation and may not always convey the intended meaning accurately.
Contact
Contact Me if you have any questions or you want to use thid dataset
Note
Original work done by SentenceTransformers
Citation
If you use the Arabic Matryoshka Embeddings Dataset, please cite it as follows:
@misc{nacar2024enhancingsemanticsimilarityunderstanding,
title={Enhancing Semantic Similarity Understanding in Arabic NLP with Nested Embedding Learning},
author={Omer Nacar and Anis Koubaa},
year={2024},
eprint={2407.21139},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2407.21139},
}