Team Ai
14 results

WikiLarge

williamplacroix /wikilarge-graded-gpt2toneizer100K<n<1M0 likes404 downloads2y agoHugging Faceeilamc14 /wikilarge-clean WikiLarge Cleaned SummaryThis dataset is a cleaned and deduplicated subset of the classic WikiLarge-style sentence pairs (English Wikipedia → Simple English Wikipedia).Starting from the original alignment files (wiki.full.aner.ori.train/valid/test.{src,dst}), we constructed a Hugging Face datasets corpus, applied a set of cheap filters, and removed near-duplicates. Provenance & License: This is a derivative of Wikipedia / Simple English Wikipedia content under CC BY-SA.The… See the full description on the dataset page: https://huggingface.co/datasets/eilamc14/wikilarge-clean.text100K<n<1M0 likes66 downloads9mo agoHugging Facewaboucay /wikilarge WikiLarge HuggingFace implementation of the WikiLarge corpus for sentence simplification gathered by Zhang, Xingxing and Lapata, Mirella. /!\ I am not one of the creators of the dataset, I just needed a HF version of this dataset and uploaded it. I encourage you to read the paper introducing the dataset: Sentence Simplification with Deep Reinforcement Learning (Zhang & Lapata, EMNLP 2017) Uses This dataset can be used to train sentence simplification… See the full description on the dataset page: https://huggingface.co/datasets/waboucay/wikilarge.2 likes60 downloads3y agoHugging Facebogdancazan /wikilarge-text-simplificationtext100K<n<1M6 likes59 downloads3y agoHugging Facewilliamplacroix /wikilarge_grade12_alpacatext10K<n<100K0 likes54 downloads1y agoHugging Facewilliamplacroix /graded_wikilargetabulartext-generation1M<n<10M0 likes54 downloads1y agoHugging Face