Team Ai
Datasetpublic

eilamc14/wikilarge-clean

WikiLarge Cleaned SummaryThis dataset is a cleaned and deduplicated subset of the classic WikiLarge-style sentence pairs (English Wikipedia → Simple English Wikipedia).Starting from the original alignment files (wiki.full.aner.ori.train/valid/test.{src,dst}), we constructed a Hugging Face datasets corpus, applied a set of cheap filters, and removed near-duplicates. Provenance & License: This is a derivative of Wikipedia / Simple English Wikipedia content under CC BY-SA.The… See the full description on the dataset page: https://huggingface.co/datasets/eilamc14/wikilarge-clean.

sourceHugging Facecc-by-sa-4.0updated 9mo agoView on Hugging Face
0likes66downloads
Dataset Card

WikiLarge Cleaned

Summary This dataset is a cleaned and deduplicated subset of the classic WikiLarge-style sentence pairs (English Wikipedia → Simple English Wikipedia). Starting from the original alignment files (wiki.full.aner.ori.train/valid/test.{src,dst}), we constructed a Hugging Face datasets corpus, applied a set of cheap filters, and removed near-duplicates.

Provenance & License: This is a derivative of Wikipedia / Simple English Wikipedia content under CC BY-SA. The derivative is distributed under CC BY-SA 4.0 as required by the ShareAlike clause. Please ensure proper attribution (TASL: Title, Author(s), Source, License) and preserve the same license for downstream derivatives.
  • —Original alignment resource (WikiLarge): [DRESS GitHub](https://github.com/XingxingZhang/dress)
  • —This dataset (code & scripts) can be found in the repository: [Simplify This Github](https://github.com/eilamc14/Simplify-This)

Data Schema

  • —Fields
  • —source (str) – original (complex) sentence from English Wikipedia
  • —target (str) – simplified sentence from Simple English Wikipedia
  • —Splits
  • —train, validation, test (derived from the original train/valid/test files)

Construction & Cleaning

Starting point We downloaded the original WikiLarge files from the official GitHub and created a standard HF DatasetDict:

  • —Files used: wiki.full.aner.ori.train.src, wiki.full.aner.ori.train.dst, and analogous valid / test.

Cheap filters (single batched pass) — thresholds used:

  • —Length constraints
  • —Source tokens: MIN_SRC_TOKENS = 4, MAX_SRC_TOKENS = 256
  • —Target tokens: MIN_TGT_TOKENS = 2, MAX_TGT_TOKENS = 256
  • —Compression ratio (simple/complex): keep if 0.40 ≤ CR ≤ 0.95
  • —Near-identity cutoff (lexical Jaccard): keep if Jaccard(source, target) < 0.98

Implementation notes (as used here):

  • —Tokenization for checks is whitespace split (str.split()).
  • —Jaccard is computed on token sets, i.e., lexical overlap.
  • —We record a per-example first-fail reason among: length → compression → near_identity.

Deduplication strategy. We deduplicate (source, target) pairs using a stable hash key:

  • —Normalize each text by lowercasing and collapsing whitespace.
  • —Compute md5(src_norm) + "#" + md5(tgt_norm); keep the first occurrence and drop repeats.

Cleaning Results & Statistics

Sizes

  • —Train: [296402 → 123862]
  • —Validation: [992 → 417]
  • —Test: [359 → 121]

Removal breakdown (exclusive, first-fail)

  • —Train
  • —length fails: [15764] (5.32%)
  • —compression fails: [156027] (57.84%)
  • —near_identity fails: [633] (0.93%)
  • —dedup: [116] (0.04%)
  • —before → after
  • —Mean CR: [0.88 → 0.70]
  • —Mean source length: [25.17 → 26.49]
  • —Mean target length: [18.51 → 18.29]
  • —Validation
  • —length fails: [57] (5.75%)
  • —compression fails: [572] (57.66%)
  • —near_identity fails: [16] (1.61%)
  • —dedup: [0] (0.00%)
  • —Test
  • —length fails: [0] (0.00%)
  • —compression fails: [237] (66.02%)
  • —near_identity fails: [55] (15.32%)
  • —dedup: [0] (0.00%)

Intended Use

  • —Primary task: Text simplification (English).
  • —Recommended splits: Use validation for model selection and test for final reporting.
  • —Caveats: Original sentence alignments were automatically mined; occasional noise or content drift may exist even after this cleaning.

Usage

python
from datasets import load_dataset

ds = load_dataset("eilamc14/wikilarge-clean")
print(ds)
print(ds["train"][0])  # {"source": "...", "target": "..."}

Attribution & License

  • —Provenance: Derived from English Wikipedia and Simple English Wikipedia content.
  • —License: CC BY-SA 4.0 (ShareAlike applies). You must provide attribution (TASL: Title, Author(s), Source, License) and distribute derivatives under the same license.

Reproducibility

Filters and dedup are implemented as described above (length, CR, Jaccard, and hash-based pair dedup). Implementation reference: [dataset_cleaning.ipynb](https://github.com/eilamc14/Simplify-This/dataset_cleaning.ipynb)


Citation

Please cite both the widely used WikiLarge reference and this cleaned subset:

Original alignment (WikiLarge): Zhang, X., & Lapata, M. (2017). Sentence Simplification with Deep Reinforcement Learning. In EMNLP 2017 (pp. 584–594). Association for Computational Linguistics. doi:10.18653/v1/D17-1062

This dataset (cleaned subset): Cohen, E., Bul, I., Inbar, D., & Loewenbach, O. (2026). Simplify-This: A Comparative Analysis of Prompt-Based and Fine-Tuned LLMs. arXiv. arxiv.org/abs/2601.05794

WikiLarge BibTeX:

@inproceedings{zhang-lapata-2017-sentence, title = {Sentence Simplification with Deep Reinforcement Learning}, author = {Zhang, Xingxing and Lapata, Mirella}, booktitle = {Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing}, pages = {584--594}, year = {2017}, address = {Copenhagen, Denmark}, publisher = {Association for Computational Linguistics}, doi = {10.18653/v1/D17-1062}, url = {https://aclanthology.org/D17-1062/} }

Wikilarge cleaned BibTeX:

@misc{simplifythis2025, author = {Cohen, Eilam and others}, title = {Simplify-This: A Comparative Analysis of Prompt-Based and Fine-Tuned LLMs}, year = {2025}, howpublished = {\url{https://github.com/eilamc14/Simplify-This}}, note = {GitHub repository}, urldate = {2025-09-30} }