datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
amiGigaSpeech is an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality
labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised
and unsupervised training. Around 40,000 hours of transcribed audio is first collected from audiobooks, podcasts
and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science,
sports, etc. A new forced alignment and segmentation pipeline is proposed to create sentence segments suitable
for speech recognition training, and to filter out segments with low-quality transcription. For system training,
GigaSpeech provides five subsets of different sizes, 10h, 250h, 1000h, 2500h, and 10000h.
For our 10,000-hour XL training subset, we cap the word error rate at 4% during the filtering/validation stage,
and for all our other smaller training subsets, we cap it at 0%. The DEV and TEST evaluation sets, on the other hand,
are re-processed by professional human transcribers to ensure high transcription quality.seq2seq-mixed-pretraining-SmolLM2seq2seq-glue
Dataset Card for "seq2seq-glue"
More Information needed
ami-ihm-kaldi-processedseq2seq-cnndm
Dataset Card for "seq2seq-cnndm"
More Information needed
seq2seq-mnli
Dataset Card for "seq2seq-mnli"
More Information needed
seq2seq-stsb
Dataset Card for "seq2seq-stsb"
More Information needed
esnli-seq2seqseq2seq-qnli
Dataset Card for "seq2seq-qnli"
More Information needed
seq2seq-qqp
Dataset Card for "seq2seq-qqp"
More Information needed
seq2seq-sst2
Dataset Card for "seq2seq-sst2"
More Information needed
seq2seq-mrpc
Dataset Card for "seq2seq-mrpc"
More Information needed
seq2seq-cnndm-tokenized
Dataset Card for "seq2seq-cnndm-tokenized"
More Information needed
phoner_seq2seq
Dataset Card for "phoner_seq2seq"
More Information needed
seq2seq-rte
Dataset Card for "seq2seq-rte"
More Information needed
seq2seq-cola
Dataset Card for "seq2seq-cola"
More Information needed
DST_MultiWOZ_seq2seqseq2seq-squad
Dataset Card for "seq2seq-squad"
More Information needed
dna2aa-eb502-seq2seq
DNA → Amino Acid Translation Dataset (EB502 Submission)
Description
이 데이터셋은 무작위로 생성된 DNA 시퀀스(A, C, G, T)와 표준 유전 코드를 사용한 해당 아미노산 번역 쌍을 포함합니다. 각 DNA 서열은 ATG로 시작하여 정지 코돈으로 끝납니다.
Dataset Structure
필드: src (DNA), tgt (아미노산)
분할: train (240개), validation (60개), test (60개)
hovh_tumanyan_seq2seq
📝 Seq2Seq Dataset: Hovhannes Tumanyan's Poems
This dataset contains sentence pairs extracted from the poetic works of Hovhannes Tumanyan, one of the most celebrated Armenian poets. It is formatted for training sequence-to-sequence (Seq2Seq) models for tasks such as:
Text generation
Dialogue modeling
Style imitation
📂 Dataset Structure
The dataset is provided as a .csv file with two columns:
input_sentence
target_sentence
Line 1 of poem
Line 2 of poem
Line… See the full description on the dataset page: https://huggingface.co/datasets/EdUarD0110/hovh_tumanyan_seq2seq.seq2seq-squad-tokenized
Dataset Card for "seq2seq-squad-tokenized"
More Information needed
en-Pile-NER-seq2seq_formatsquad_seq2seqreddit_one_ups_seq2seq_2014
Dataset Card for reddit_one_ups_seq2seq_2014
Dataset Summary
Reddit 'one-ups' or 'clapbacks' - replies which scored higher than the original comments.
This dataset chose freeform replies, which did not follow repetitive meme replies. The IAmA subreddit was excluded to avoid an issue where their answers frequently score higher than questions.
For commentary on predictions with a previous version of the dataset, see… See the full description on the dataset page: https://huggingface.co/datasets/georeactor/reddit_one_ups_seq2seq_2014.budget-seq2seq
Dataset Card for "budget-seq2seq-xml"
More Information needed
diffusiondb-seq2seq
Dataset Card for "diffusiondb-seq2seq"
More Information needed
propbank_srl_seq2seqbudget-seq2seq-json
Dataset Card for "budget-seq2seq-json"
More Information needed
datasets_statslsoie_seq2seq
