datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Long-Instruction-with-Paraphrasing
🔥 Updates
[2024.6.4] Add a slim version. The sample number is reduced from about 20k to 10k.
[2024.5.28]
The data format is converted from "chatml" to "messages", which is more convenient to use tokenizer.apply_chat_template. The old version has been moved to "legacy" branch.
The version without "Original text paraphrasing" is added.
📊 Long Context Instruction-tuning dataset with "Original text paraphrasing"
Paper
Github
consist of multiple tasks
Chinese and… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/Long-Instruction-with-Paraphrasing.task045_miscellaneous_sentence_paraphrasing
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task045_miscellaneous_sentence_paraphrasing
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task045_miscellaneous_sentence_paraphrasing.task1288_glue_mrpc_paraphrasing
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1288_glue_mrpc_paraphrasing
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1288_glue_mrpc_paraphrasing.paraphrasing-preferences-orpo-dpo
Paraphrasing Preference Dataset
A preference dataset for training paraphrase models via DPO, RLHF, or ORPO. Each example contains a source text, a task-specific prompt, and a chosen/rejected paraphrase pair ranked by a composite quality score.
Dataset Summary
Train
Val
Total
Examples
852
95
947
Sources: Quora questions (571), SQuAD 2.0 sentences (218), CNN News sentences (158). The val split is stratified by excellent_in, category, and binned total_delta.… See the full description on the dataset page: https://huggingface.co/datasets/alecccdd/paraphrasing-preferences-orpo-dpo.task177_para-nmt_paraphrasing
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task177_para-nmt_paraphrasing
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task177_para-nmt_paraphrasing.Identification-of-paraphrasing
🇰🇿 Identification of Paraphrasing in Kazakh Context
Dataset Summary
Identification of Paraphrasing in Kazakh Context is a targeted dataset designed to train Large Language Models (LLMs) and embeddings to detect semantic equivalence between two distinct Kazakh texts.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
2,000
Total Words (approx.)
184,465
Avg. Words per Sample
92
Word Count… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Identification-of-paraphrasing.
