Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Congi-libya /bayan-simplification-corpus Synthetic simplification data card — v1 Versions. main and the v1 tag hold v1 (Gemma 4 31B generator, Qwen3.8 27B judge). The previous DeepSeek release is kept at the v0 tag: load_dataset("Congi-libya/bayan-simplification-corpus", revision="v0"). A full v0 vs v1 comparison is in Cogni-Libya/Bayan#45. Paths under scripts/ and docs/ refer to the Bayan repository; paths under data/processed/ are the project's working files and are not part of this dataset. Version 1 of… See the full description on the dataset page: https://huggingface.co/datasets/Congi-libya/bayan-simplification-corpus.tabulartext-generation10K<n<100K3 likes238 downloads11d agoHugging Face02r1char9 /simplification-datasetДанный dataset был собран из корпуса "RuSimpleSentEval" (https://github.com/dialogue-evaluation/RuSimpleSentEval), а также "RuAdapt" (https://github.com/Digital-Pushkin-Lab/RuAdapt) для задачи упрощения текста (text simplification). from datasets import load_dataset data_files = {'train':"train.csv",'test':"test.csv"} dataset = load_dataset("r1char9/simplification", data_files=data_files) train_df = dataset['train'].to_pandas() test_df = dataset['test'].to_pandas() text1K<n<10K0 likes116 downloads1mo agoHugging Face03hasankursun /age-specific-text-simplification Age-Specific Text Simplification Dataset Dataset Description This dataset contains complex texts simplified into age-appropriate versions for children aged 3, 4, and 5 years old. Each original text has been professionally adapted to match the cognitive development, vocabulary, and comprehension abilities of each specific age group. Dataset Summary Total Examples: 17,177 Training Split: 15,459 examples Validation Split: 1,718 examples Languages:… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/age-specific-text-simplification.tabulartext-generation10K<n<100K3 likes114 downloads1y agoHugging Face04NecroMOnk /bfs-optimal-symbolic-algebra-simplification-trajectories ISRE v7 BFS Trajectories This dataset contains BFS-optimal symbolic algebra simplification trajectories for the ISRE research project. Each row is one trajectory. Nested fields are stored as JSON strings to preserve the original AST and step structure without lossy flattening. What This Dataset Is This is not a natural-language instruction dataset. It is a symbolic algebra policy-learning dataset. Each trajectory starts from a deliberately scrambled algebraic… See the full description on the dataset page: https://huggingface.co/datasets/NecroMOnk/bfs-optimal-symbolic-algebra-simplification-trajectories.tabulartext-classification10K<n<100K0 likes102 downloads4mo agoHugging Face05Lots-of-LoRAs /task934_turk_simplification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task934_turk_simplification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task934_turk_simplification.texttext-generation1K<n<10K0 likes76 downloads2y agoHugging Face06vulturuldemare /Estonian-Text-Simplification Estonian Text Simplification This repository contains resources and models for Estonian text simplification, including datasets and pre-trained models. Files Dataset Files simplification_training_set.json: A dataset used to fine-tune LLaMA 3.1 for text simplification. src: The source of the data. original: The original sentence to be simplified. simpl_lex: A lexical simplification (may be empty). simpl_final: The final simplified sentence.… See the full description on the dataset page: https://huggingface.co/datasets/vulturuldemare/Estonian-Text-Simplification.text10K<n<100K0 likes75 downloads2y agoHugging Face07BramVanroy /wiki_simplifications_dutch_dedup_splitThis is a variant of the original dataset. It was shuffled (seed=42); Deduplicated on rows (96,613 rows removed); Split into train, validation and test sets (the latter have 8192 samples each) Reproduction from datasets import load_dataset, Dataset, DatasetDict ds = load_dataset("UWV/Leesplank_NL_wikipedia_simplifications", split="train") ds = ds.shuffle(seed=42) print("original", ds) df = ds.to_pandas() df = df.drop_duplicates().reset_index() ds = Dataset.from_pandas(df)… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/wiki_simplifications_dutch_dedup_split.text1M<n<10M1 likes63 downloads3y agoHugging Face08melll-uff /legal-simplification-pt Dataset Structure Configuration: random_split train_random validation_random test_random Configuration: by_source acordaos_tcu stf_decisions stf_votes tjsp trf5 ulysses_tesemo tabular100K<n<1M1 likes60 downloads6mo agoHugging Face09BramVanroy /chatgpt-dutch-simplification Dataset Card for ChatGPT Dutch Simplification Dataset Summary Created in light of a master thesis by Charlotte Van de Velde as part of the Master of Science in Artificial Intelligence at KU Leuven. Charlotte is supervised by Vincent Vandeghinste and Bram Vanroy. The dataset contains Dutch source sentences and aligned simplified sentences, generated with ChatGPT. All splits combined, the dataset consists of 1267 entries. Charlotte used gpt-3.5-turbo with the following… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/chatgpt-dutch-simplification.text1K<n<10K5 likes59 downloads3y agoHugging Face10bogdancazan /wikilarge-text-simplificationtext100K<n<1M6 likes59 downloads3y agoHugging Face11CATIE-AQ /bisect_fr_prompt_textual_simplification bisect_fr_prompt_textual_simplification Summary bisect_fr_prompt_textual_simplification is a subset of the Dataset of French Prompts (DFP).It contains 9,889,420 rows that can be used for a textual simplification task.The original data (without prompts) comes from the dataset BiSECT by Kim et al. where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/bisect_fr_prompt_textual_simplification.textsummarization1M<n<10M0 likes53 downloads1y agoHugging Face12UWV /Leesplank_NL_wikipedia_simplificationsThe set contains 2.87M pragraphs of prompt/result combinations, where the prompt is a paragraph from Dutch Wikipedia and the result is a simplified text, which could include more than one paragraph. This dataset was created by UWV, as a part of project "Leesplank", an effort to generate datasets that are ethically and legally sound. The basis of this dataset was the wikipedia extract as a part of Gigacorpus (http://gigacorpus.nl/). The lines were fed one by one into GPT 4 1106 preview, where… See the full description on the dataset page: https://huggingface.co/datasets/UWV/Leesplank_NL_wikipedia_simplifications.texttext-generation1M<n<10M6 likes47 downloads3y agoHugging Face13UWV /Leesplank_NL_wikipedia_simplifications_preprocessed Dataset Card for Dataset Name A synthetic dataset made by simplifying Dutch Wikipedia entries. This is a processed version of an earlier publication. Dataset Details Dataset Description Leesplank_nl_wikipedia_simplifications_preprocessed is based on https://huggingface.co/datasets/BramVanroy/wiki_simplifications_dutch_dedup_split, which in itself is based on our own dataset (https://huggingface.co/datasets/UWV/Leesplank_NL_wikipedia_simplifications), but… See the full description on the dataset page: https://huggingface.co/datasets/UWV/Leesplank_NL_wikipedia_simplifications_preprocessed.text1M<n<10M3 likes45 downloads2y agoHugging Face14Sadou /medical-reports-simplification-dataset 🏥 Medical Reports Simplification Dataset 📋 Description Dataset créé avec Gemini 2.5 Pro (Preview) pour entraîner des modèles à simplifier les rapports médicaux complexes en explications compréhensibles pour les patients. 🎯 Objectif : Démocratiser l'accès à l'information médicale en rendant les rapports techniques accessibles au grand public. 🔧 Génération du Dataset Génération : Gemini 2.5 Pro (Preview) Validation : Contrôle qualité automatisé… See the full description on the dataset page: https://huggingface.co/datasets/Sadou/medical-reports-simplification-dataset.texttext-generationn<1K0 likes39 downloads1y agoHugging Face15Lots-of-LoRAs /task111_asset_sentence_simplification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task111_asset_sentence_simplification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task111_asset_sentence_simplification.texttext-generation1K<n<10K0 likes35 downloads2y agoHugging Face16lighteval /med_paragraph_simplificationtext1K<n<10K0 likes31 downloads1y agoHugging Face17khaledmahmoud /spanish_simplification_20k Dataset Card: spanish_simplification_20k This dataset provides 20K samples for Spanish simplification across a diverse range of topics, designed for people with cognitive challenges. The samples vary in length, from short to long. This dataset can supplement a larger training corpus for fine-tuning small language models, such as Gemma 4B, to simplify complex Spanish into simpler Spanish or translate English directly into simplified Spanish. Dataset Contributors… See the full description on the dataset page: https://huggingface.co/datasets/khaledmahmoud/spanish_simplification_20k.tabular10K<n<100K0 likes26 downloads5mo agoHugging Face18rpisano /nemotron-cc-atomic-simplification-gemma4-31b nemotron-cc atomic-statement simplification (Gemma 4 31B-it) 2,000,000 records: source text from nvidia/nemotron-cc-v2.1 (High-Quality-Synthetic split) rewritten by google/gemma-4-31B-it into a sequence of atomic, Subject-Verb-Object statements. Generated with vLLM 0.22.1 in-process batch inference (see src/generate/run.py in the producing repo), TP=4, max_model_len=16384, max_tokens=8192, prompts filtered to <=8192 templated tokens. Fields id: original… See the full description on the dataset page: https://huggingface.co/datasets/rpisano/nemotron-cc-atomic-simplification-gemma4-31b.texttext-generation1M<n<10M0 likes26 downloads1mo agoHugging Face19athugodage /legal_simplification_corpus Dataset Card for "legal_simplification_corpus" More Information needed text1K<n<10K0 likes24 downloads4y agoHugging Face20raja20221020 /english-text-simplification-for-finetuningtext10K<n<100K0 likes21 downloads1y agoHugging Face21bogdancazan /news-not-not-ela-text-simplificationtext100K<n<1M3 likes20 downloads3y agoHugging Face22bogdancazan /biendata_text_simplificationtext10K<n<100K0 likes18 downloads3y agoHugging Face23supergoose /flan_combined_task934_turk_simplificationtext1K<n<10K0 likes17 downloads2y agoHugging Face24Karthikrv /adaption-legal-clause-simplification This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-legal_clause_simplification This dataset contains pairs of formal legal or policy clauses and their simplified, plain-language equivalents. The prompts feature complex contractual language regarding payments, confidentiality, and benefits, while the completions provide clear, concise restatements. It is designed for training models to translate legalese into accessible text.… See the full description on the dataset page: https://huggingface.co/datasets/Karthikrv/adaption-legal-clause-simplification.text1K<n<10K0 likes13 downloads4mo agoHugging Face25Karthikrv /adaption-legal-text-simplification This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-legal_text_simplification This dataset contains pairs of formal legal or institutional text prompts and their corresponding simplified, plain-language completions. The content focuses on translating complex terms regarding consumer responsibilities, examination policies, and liability into clear, everyday English. It is designed for training models to perform legal text… See the full description on the dataset page: https://huggingface.co/datasets/Karthikrv/adaption-legal-text-simplification.text1K<n<10K0 likes11 downloads4mo agoHugging Face26igornishka /dutch-municipal-sentence-simplificationtext1K<n<10K0 likes10 downloads3y agoHugging Face27supergoose /flan_combined_task111_asset_sentence_simplificationtext1K<n<10K0 likes10 downloads2y agoHugging Face28Nechba /wikilarge-text-simplificationtext100K<n<1M0 likes10 downloads6mo agoHugging Face29raja20221020 /hindi-text-simplification-for-finetuningtext10K<n<100K0 likes8 downloads1y agoHugging Face30marcov /asset_simplification_promptsourcetext10K<n<100K0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.