Team Ai
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Vibels /parallel-corpus_geo_eng_geo-translatedData gathered from https://enkacorpus.iliauni.edu.ge/, active research corpus which provides georgian english translated sentences. Afterwards english part was translated using Google Cloud Translation API v3 and appended as eng_geo column. Intentional usage: research and learning Intentional task categories: Sentence Similarity, Translation texttranslation100K<n<1M0 likes26 downloads1y agoHugging Face02joanitolopo /kupangmalay-parallelcorpustext10K<n<100K0 likes24 downloads2y agoHugging Face03Arseniy-Polyakov /parallel_corpus_russian_rsl_glossestext1K<n<10K0 likes23 downloads7mo agoHugging Face04Thermostatic /parallel_corpus_europarl_english_spanish Dataset Card for Dataset Name A massive parallel corpus of English-Spanish pairs. It hasn't a specified license, but there doesn't seem to be any copyrighted material in the corpus. Dataset Details Dataset Description Curated by: Philipp Koehn Funded by [optional]: In part funded by the European Commission (7th Framework Programme). Shared by [optional]: [More Information Needed] Language(s) (NLP): English & Spanish License: Not specified.… See the full description on the dataset page: https://huggingface.co/datasets/Thermostatic/parallel_corpus_europarl_english_spanish.texttranslation1M<n<10M1 likes18 downloads3y agoHugging Face05navinaananthan /Dhivehi-English-ParallelCorpustext100K<n<1M1 likes16 downloads3y agoHugging Face06Thermostatic /parallel_corpus_webcrawl_english_spanish_1 Dataset Card for Dataset Name This parallel corpus dataset contains about 21k rows of parallel English and Spanish texts obtained by crawling different websites. It has been filtered strictly. Dataset Details Dataset Description This is a parallel corpus of bilingual texts crawled from multilingual websites, which contains 21, 005 TUs. A strict validation process has been followed, which resulted in discarding: TUs from crawled websites that do not comply… See the full description on the dataset page: https://huggingface.co/datasets/Thermostatic/parallel_corpus_webcrawl_english_spanish_1.texttranslation10K<n<100K1 likes11 downloads3y agoHugging Face07ashuChufamo /parallel-corpus_en-amtexttranslation10K<n<100K0 likes9 downloads2y agoHugging Face08ChengSong /parallel-corpus-balanced-en-pttext10K<n<100K0 likes9 downloads11mo agoHugging Face09TenzinKhorloVIT /ParallelCorpustext100K<n<1M0 likes4 downloads3y agoHugging Face10ChengSong /parallel-corpus-total-en-pttext1M<n<10M0 likes4 downloads11mo agoHugging Face11ChengSong /parallel-corpus-total-pt-zhtext1M<n<10M0 likes3 downloads11mo agoHugging Face12ChengSong /parallel-corpus-balanced-pt-zhtext10K<n<100K0 likes2 downloads11mo agoHugging Face13ltsab1618033988 /parallel-corpus-en-kagatedtext100K<n<1M1 likes1 downloads3y agoHugging Face14Mobiusi /Parallel-Corpus-Dataset-Of-Land-Use-Zoning-And-Development-Control-Texts Parallel Corpus Dataset of Land-Use Zoning and Development Control Texts This corpus pairs source texts with target-language translations of regulatory clauses from land use planning, including zoning controls, permitted uses, development restrictions, and planning controls. It captures formal regulatory phrasing and cross-language equivalents for planning terminology. Fields for translation instructions, language information, clause categories, and aligned terms support… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Parallel-Corpus-Dataset-Of-Land-Use-Zoning-And-Development-Control-Texts.texttext-classificationn<1K0 likes13h agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.