Simplification
simplificationcochrane-simplificationThis dataset measures the ability for a model to simplify paragraphs of medical text through the omission non-salient information and simplification of medical jargon.bayan-simplification-corpus
Synthetic simplification data card — v1
Versions. main and the v1 tag hold v1 (Gemma 4 31B generator, Qwen3.8 27B judge). The previous
DeepSeek release is kept at the v0 tag: load_dataset("Congi-libya/bayan-simplification-corpus", revision="v0").
A full v0 vs v1 comparison is in Cogni-Libya/Bayan#45.
Paths under scripts/ and docs/ refer to the Bayan repository;
paths under data/processed/ are the project's working files and are not part of this dataset.
Version 1 of… See the full description on the dataset page: https://huggingface.co/datasets/Congi-libya/bayan-simplification-corpus.unsat_2024_batch3000_8gpu_no_simplification_lowest7_multiassignsimplification-datasetДанный dataset был собран из корпуса "RuSimpleSentEval" (https://github.com/dialogue-evaluation/RuSimpleSentEval), а также "RuAdapt" (https://github.com/Digital-Pushkin-Lab/RuAdapt) для задачи упрощения текста (text simplification).
from datasets import load_dataset
data_files = {'train':"train.csv",'test':"test.csv"}
dataset = load_dataset("r1char9/simplification", data_files=data_files)
train_df = dataset['train'].to_pandas()
test_df = dataset['test'].to_pandas()
age-specific-text-simplification
Age-Specific Text Simplification Dataset
Dataset Description
This dataset contains complex texts simplified into age-appropriate versions for children aged 3, 4, and 5 years old. Each original text has been professionally adapted to match the cognitive development, vocabulary, and comprehension abilities of each specific age group.
Dataset Summary
Total Examples: 17,177
Training Split: 15,459 examples
Validation Split: 1,718 examples
Languages:… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/age-specific-text-simplification.
