Team Ai
Datasetpublic

ilsp/scipar_parallel_docs

SciPar Parallel Documents Dataset Description This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts. In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories. This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences. To do this, we kept only the parallel… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.

sourceHugging Facecc-by-nc-sa-4.0updated 3y agoView on Hugging Face
2likes152downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

Team Ai shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
ilsp/scipar_parallel_docs · Team Ai