Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01speech-seq2seq /amiGigaSpeech is an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and unsupervised training. Around 40,000 hours of transcribed audio is first collected from audiobooks, podcasts and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science, sports, etc. A new forced alignment and segmentation pipeline is proposed to create sentence segments suitable for speech recognition training, and to filter out segments with low-quality transcription. For system training, GigaSpeech provides five subsets of different sizes, 10h, 250h, 1000h, 2500h, and 10000h. For our 10,000-hour XL training subset, we cap the word error rate at 4% during the filtering/validation stage, and for all our other smaller training subsets, we cap it at 0%. The DEV and TEST evaluation sets, on the other hand, are re-processed by professional human transcribers to ensure high transcription quality.0 likes2.7k downloads4y agoHugging Face02aklein4 /seq2seq-mixed-pretraining-SmolLM2tabular100M<n<1B1 likes661 downloads9mo agoHugging Face03closji /seq2seq-glue Dataset Card for "seq2seq-glue" More Information needed text1M<n<10M0 likes214 downloads3y agoHugging Face04speech-seq2seq /ami-ihm-kaldi-processed0 likes129 downloads4y agoHugging Face05closji /seq2seq-cnndm Dataset Card for "seq2seq-cnndm" More Information needed text100K<n<1M0 likes95 downloads3y agoHugging Face06closji /seq2seq-mnli Dataset Card for "seq2seq-mnli" More Information needed text100K<n<1M0 likes57 downloads3y agoHugging Face07closji /seq2seq-stsb Dataset Card for "seq2seq-stsb" More Information needed text1K<n<10K0 likes54 downloads3y agoHugging Face08ErfanMoosaviMonazzah /esnli-seq2seqtext100K<n<1M0 likes52 downloads3y agoHugging Face09closji /seq2seq-qnli Dataset Card for "seq2seq-qnli" More Information needed text100K<n<1M0 likes50 downloads3y agoHugging Face10closji /seq2seq-qqp Dataset Card for "seq2seq-qqp" More Information needed text100K<n<1M0 likes46 downloads3y agoHugging Face11closji /seq2seq-sst2 Dataset Card for "seq2seq-sst2" More Information needed text10K<n<100K0 likes42 downloads3y agoHugging Face12closji /seq2seq-mrpc Dataset Card for "seq2seq-mrpc" More Information needed text1K<n<10K0 likes40 downloads3y agoHugging Face13closji /seq2seq-cnndm-tokenized Dataset Card for "seq2seq-cnndm-tokenized" More Information needed 100K<n<1M0 likes35 downloads3y agoHugging Face14Yuhthe /phoner_seq2seq Dataset Card for "phoner_seq2seq" More Information needed text10K<n<100K0 likes34 downloads3y agoHugging Face15closji /seq2seq-rte Dataset Card for "seq2seq-rte" More Information needed text1K<n<10K0 likes31 downloads3y agoHugging Face16closji /seq2seq-cola Dataset Card for "seq2seq-cola" More Information needed text10K<n<100K0 likes28 downloads3y agoHugging Face17leftattention /DST_MultiWOZ_seq2seq0 likes28 downloads3y agoHugging Face18closji /seq2seq-squad Dataset Card for "seq2seq-squad" More Information needed text10K<n<100K0 likes27 downloads3y agoHugging Face19seoyoungyoon /dna2aa-eb502-seq2seq DNA → Amino Acid Translation Dataset (EB502 Submission) Description 이 데이터셋은 무작위로 생성된 DNA 시퀀스(A, C, G, T)와 표준 유전 코드를 사용한 해당 아미노산 번역 쌍을 포함합니다. 각 DNA 서열은 ATG로 시작하여 정지 코돈으로 끝납니다. Dataset Structure 필드: src (DNA), tgt (아미노산) 분할: train (240개), validation (60개), test (60개) textn<1K0 likes24 downloads11mo agoHugging Face20EdUarD0110 /hovh_tumanyan_seq2seq 📝 Seq2Seq Dataset: Hovhannes Tumanyan's Poems This dataset contains sentence pairs extracted from the poetic works of Hovhannes Tumanyan, one of the most celebrated Armenian poets. It is formatted for training sequence-to-sequence (Seq2Seq) models for tasks such as: Text generation Dialogue modeling Style imitation 📂 Dataset Structure The dataset is provided as a .csv file with two columns: input_sentence target_sentence Line 1 of poem Line 2 of poem Line… See the full description on the dataset page: https://huggingface.co/datasets/EdUarD0110/hovh_tumanyan_seq2seq.text10K<n<100K1 likes21 downloads1y agoHugging Face21closji /seq2seq-squad-tokenized Dataset Card for "seq2seq-squad-tokenized" More Information needed 10K<n<100K0 likes19 downloads3y agoHugging Face22vietnqw /en-Pile-NER-seq2seq_formattext100K<n<1M0 likes19 downloads3y agoHugging Face23yuvalkirstain /squad_seq2seqtext10K<n<100K0 likes18 downloads5y agoHugging Face24georeactor /reddit_one_ups_seq2seq_2014 Dataset Card for reddit_one_ups_seq2seq_2014 Dataset Summary Reddit 'one-ups' or 'clapbacks' - replies which scored higher than the original comments. This dataset chose freeform replies, which did not follow repetitive meme replies. The IAmA subreddit was excluded to avoid an issue where their answers frequently score higher than questions. For commentary on predictions with a previous version of the dataset, see… See the full description on the dataset page: https://huggingface.co/datasets/georeactor/reddit_one_ups_seq2seq_2014.tabular10K<n<100K2 likes13 downloads4y agoHugging Face25napatswift /budget-seq2seq Dataset Card for "budget-seq2seq-xml" More Information needed text10K<n<100K0 likes13 downloads3y agoHugging Face26roborovski /diffusiondb-seq2seq Dataset Card for "diffusiondb-seq2seq" More Information needed text10K<n<100K0 likes11 downloads3y agoHugging Face27cu-kairos /propbank_srl_seq2seqtext10K<n<100K0 likes11 downloads2y agoHugging Face28napatswift /budget-seq2seq-json Dataset Card for "budget-seq2seq-json" More Information needed text10K<n<100K0 likes8 downloads3y agoHugging Face29speech-seq2seq /datasets_stats0 likes7 downloads4y agoHugging Face30Thanmay /lsoie_seq2seqtext10K<n<100K0 likes7 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.