datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kupangmalay-parallelcorpusparallel-corpus_geo_eng_geo-translatedData gathered from https://enkacorpus.iliauni.edu.ge/, active research corpus which provides georgian english translated sentences. Afterwards english part was translated using Google Cloud Translation API v3 and appended as eng_geo column.
Intentional usage: research and learning
Intentional task categories: Sentence Similarity, Translation
parallel_corpus_russian_rsl_glossesparallel_corpus_europarl_english_spanish
Dataset Card for Dataset Name
A massive parallel corpus of English-Spanish pairs. It hasn't a specified license, but there doesn't seem to be any copyrighted material in the corpus.
Dataset Details
Dataset Description
Curated by: Philipp Koehn
Funded by [optional]: In part funded by the European Commission (7th Framework Programme).
Shared by [optional]: [More Information Needed]
Language(s) (NLP): English & Spanish
License: Not specified.… See the full description on the dataset page: https://huggingface.co/datasets/Thermostatic/parallel_corpus_europarl_english_spanish.Dhivehi-English-ParallelCorpusParallel_corpus_of_PhD_dissertation_abstracts
[!NOTE]
Dataset origin: https://inventory.clarin.gr/corpus/886
Description
The JRC-Acquis subcorpus EL-FR (Hunalign aligned-XML) is a parallel subcorpus for French and Greek, subset of the JRC-Acquis Multilingual Parallel Corpus.
Citation
Parallel corpus of PhD dissertation abstracts (2020). [Dataset (Text corpus)]. CLARIN:EL. http://hdl.handle.net/11500/CLARIN-EL-0000-0000-682E-9
parallel_corpus_webcrawl_english_spanish_1
Dataset Card for Dataset Name
This parallel corpus dataset contains about 21k rows of parallel English and Spanish texts obtained by crawling different websites. It has been filtered strictly.
Dataset Details
Dataset Description
This is a parallel corpus of bilingual texts crawled from multilingual websites, which contains 21, 005 TUs. A strict validation process has been followed, which resulted in discarding:
TUs from crawled websites that do not comply… See the full description on the dataset page: https://huggingface.co/datasets/Thermostatic/parallel_corpus_webcrawl_english_spanish_1.parallel-corpus-balanced-en-ptparallel-corpus_en-amParallelCorpusParallel_corpus_newsletters_IFG_FR-GR
[!NOTE]
Dataset origin: https://inventory.clarin.gr/corpus/507
Description
Parallel corpus of IFT's newsletters. Source language: FR, target language: EL. Bilingual corpus compiled by the clarin:el team AUTH, including newsletters of Institut Français de Grèce.
Citation
Parallel corpus newsletters IFG FR-GR (2021, March 22). [Dataset (Text corpus)]. CLARIN:EL. http://hdl.handle.net/11500/CLARIN-EL-0000-0000-68E2-C
parallel-corpus-total-en-ptparallel-corpus-total-pt-zhparallel-corpus-balanced-pt-zhparallel-corpus-en-ka
