datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
parallel-corpus_geo_eng_geo-translatedData gathered from https://enkacorpus.iliauni.edu.ge/, active research corpus which provides georgian english translated sentences. Afterwards english part was translated using Google Cloud Translation API v3 and appended as eng_geo column.
Intentional usage: research and learning
Intentional task categories: Sentence Similarity, Translation
kupangmalay-parallelcorpusparallel_corpus_russian_rsl_glossesparallel_corpus_europarl_english_spanish
Dataset Card for Dataset Name
A massive parallel corpus of English-Spanish pairs. It hasn't a specified license, but there doesn't seem to be any copyrighted material in the corpus.
Dataset Details
Dataset Description
Curated by: Philipp Koehn
Funded by [optional]: In part funded by the European Commission (7th Framework Programme).
Shared by [optional]: [More Information Needed]
Language(s) (NLP): English & Spanish
License: Not specified.… See the full description on the dataset page: https://huggingface.co/datasets/Thermostatic/parallel_corpus_europarl_english_spanish.Dhivehi-English-ParallelCorpusparallel_corpus_webcrawl_english_spanish_1
Dataset Card for Dataset Name
This parallel corpus dataset contains about 21k rows of parallel English and Spanish texts obtained by crawling different websites. It has been filtered strictly.
Dataset Details
Dataset Description
This is a parallel corpus of bilingual texts crawled from multilingual websites, which contains 21, 005 TUs. A strict validation process has been followed, which resulted in discarding:
TUs from crawled websites that do not comply… See the full description on the dataset page: https://huggingface.co/datasets/Thermostatic/parallel_corpus_webcrawl_english_spanish_1.parallel-corpus_en-amparallel-corpus-balanced-en-ptParallelCorpusparallel-corpus-total-en-ptparallel-corpus-total-pt-zhparallel-corpus-balanced-pt-zhparallel-corpus-en-kaParallel-Corpus-Dataset-Of-Land-Use-Zoning-And-Development-Control-Texts
Parallel Corpus Dataset of Land-Use Zoning and Development Control Texts
This corpus pairs source texts with target-language translations of regulatory clauses from land use planning, including zoning controls, permitted uses, development restrictions, and planning controls. It captures formal regulatory phrasing and cross-language equivalents for planning terminology. Fields for translation instructions, language information, clause categories, and aligned terms support… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Parallel-Corpus-Dataset-Of-Land-Use-Zoning-And-Development-Control-Texts.
