Team Ai
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Congi-libya /bayan-simplification-corpus Synthetic simplification data card — v1 Versions. main and the v1 tag hold v1 (Gemma 4 31B generator, Qwen3.8 27B judge). The previous DeepSeek release is kept at the v0 tag: load_dataset("Congi-libya/bayan-simplification-corpus", revision="v0"). A full v0 vs v1 comparison is in Cogni-Libya/Bayan#45. Paths under scripts/ and docs/ refer to the Bayan repository; paths under data/processed/ are the project's working files and are not part of this dataset. Version 1 of… See the full description on the dataset page: https://huggingface.co/datasets/Congi-libya/bayan-simplification-corpus.tabulartext-generation10K<n<100K3 likes238 downloads11d agoHugging Face02hasankursun /age-specific-text-simplification Age-Specific Text Simplification Dataset Dataset Description This dataset contains complex texts simplified into age-appropriate versions for children aged 3, 4, and 5 years old. Each original text has been professionally adapted to match the cognitive development, vocabulary, and comprehension abilities of each specific age group. Dataset Summary Total Examples: 17,177 Training Split: 15,459 examples Validation Split: 1,718 examples Languages:… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/age-specific-text-simplification.tabulartext-generation10K<n<100K3 likes114 downloads1y agoHugging Face03liliya-makhmutova /medical_texts_simplification Dataset Card for Medical texts simplification The dataset consisting of 30 triples (around 800 sentences) of the original text, human- and ChatGPT-simplified texts was created from a subset Medical Notes Classification dataset. The original dataset contains medical notes, which come from exactly one of the following five clinical domains: Gastroenterology, Neurology, Orthopedics, Radiology, and Urology. There are 1239 texts in total in the original dataset. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/liliya-makhmutova/medical_texts_simplification.text-generationn<1K4 likes80 downloads3y agoHugging Face04Lots-of-LoRAs /task934_turk_simplification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task934_turk_simplification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task934_turk_simplification.texttext-generation1K<n<10K0 likes76 downloads2y agoHugging Face05UWV /Leesplank_NL_wikipedia_simplificationsThe set contains 2.87M pragraphs of prompt/result combinations, where the prompt is a paragraph from Dutch Wikipedia and the result is a simplified text, which could include more than one paragraph. This dataset was created by UWV, as a part of project "Leesplank", an effort to generate datasets that are ethically and legally sound. The basis of this dataset was the wikipedia extract as a part of Gigacorpus (http://gigacorpus.nl/). The lines were fed one by one into GPT 4 1106 preview, where… See the full description on the dataset page: https://huggingface.co/datasets/UWV/Leesplank_NL_wikipedia_simplifications.texttext-generation1M<n<10M6 likes47 downloads3y agoHugging Face06Sadou /medical-reports-simplification-dataset 🏥 Medical Reports Simplification Dataset 📋 Description Dataset créé avec Gemini 2.5 Pro (Preview) pour entraîner des modèles à simplifier les rapports médicaux complexes en explications compréhensibles pour les patients. 🎯 Objectif : Démocratiser l'accès à l'information médicale en rendant les rapports techniques accessibles au grand public. 🔧 Génération du Dataset Génération : Gemini 2.5 Pro (Preview) Validation : Contrôle qualité automatisé… See the full description on the dataset page: https://huggingface.co/datasets/Sadou/medical-reports-simplification-dataset.texttext-generationn<1K0 likes39 downloads1y agoHugging Face07Lots-of-LoRAs /task111_asset_sentence_simplification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task111_asset_sentence_simplification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task111_asset_sentence_simplification.texttext-generation1K<n<10K0 likes35 downloads2y agoHugging Face08rpisano /nemotron-cc-atomic-simplification-gemma4-31b nemotron-cc atomic-statement simplification (Gemma 4 31B-it) 2,000,000 records: source text from nvidia/nemotron-cc-v2.1 (High-Quality-Synthetic split) rewritten by google/gemma-4-31B-it into a sequence of atomic, Subject-Verb-Object statements. Generated with vLLM 0.22.1 in-process batch inference (see src/generate/run.py in the producing repo), TP=4, max_model_len=16384, max_tokens=8192, prompts filtered to <=8192 templated tokens. Fields id: original… See the full description on the dataset page: https://huggingface.co/datasets/rpisano/nemotron-cc-atomic-simplification-gemma4-31b.texttext-generation1M<n<10M0 likes26 downloads1mo agoHugging Face09dispatchAI /arabic-text-simplification Arabic Text Simplification Dataset Complex-to-simplified Arabic text pairs for accessibility fine-tuning. Why This Matters Arabic text simplification is critical for: Low-literacy readers — 1 in 5 Arabic speakers struggle with complex text Cognitive accessibility — dyslexia, intellectual disabilities, autism Non-native speakers — Arabic learners and expatriate workers Children's content — making educational material age-appropriate Almost no Arabic accessibility… See the full description on the dataset page: https://huggingface.co/datasets/dispatchAI/arabic-text-simplification.text-generationn<1K0 likes24 downloads4mo agoHugging Face10go76dof /Fineweb_simplification_pairsgated FineWeb Simplification Pairs This dataset contains the FineWeb sentence-level simplification pairs used in our BabyLM pretraining experiments. The simplified rewrites were generated automatically with Qwen 7B. The corpus pairs original FineWeb sentences with meaning-preserving simplified rewrites generated by Qwen 7B. It was used to train models such as go76dof/wwm_curriculum_simplification_40k, where the model sees original and simplified text during masked language model… See the full description on the dataset page: https://huggingface.co/datasets/go76dof/Fineweb_simplification_pairs.text-generation100K<n<1M1 likes4 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.