Team Ai
15 results

shuffled

redmoddata /nemo23_shuffledtext10M<n<100M0 likes2.8k downloads8mo agoHugging Faceyuezhouhu /dolma3_dolmino_mix-100B-1125-tokenized-shuffled0 likes2.2k downloads4d agoHugging FaceHuggingFaceFW /fineweb_edu_100BT-shuffled FineWeb-Edu 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/fineweb_edu_100BT. Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as fineweb_edu_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into memory, shuffled… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb_edu_100BT-shuffled.tabular100M<n<1B6 likes2k downloads7mo agoHugging FaceDanielGallagherIRE /FineWeb-Edu-10B-Shuffledtext1M<n<10M0 likes2k downloads4mo agoHugging Facedanbraunai /pile-uncopyrighted-tok-shuffled Pile Uncopyrighted (Tokenized + Shuffled) Globally shuffled version of the Pile (uncopyrighted subset), tokenized into fixed-length sequences of 513 token IDs using EleutherAI/gpt-neox-20b. Each row contains a single input_ids column (513 int32 values). Documents are concatenated with EOS tokens between them, then reshaped into fixed-length sequences (following the TransformerLens approach). Sequences are then globally shuffled (seed=42) so that consecutive rows are not from the… See the full description on the dataset page: https://huggingface.co/datasets/danbraunai/pile-uncopyrighted-tok-shuffled.100M<n<1B0 likes1.5k downloads8mo agoHugging FaceMuesli1 /dclm-baseline-1.0-llama3-tokenized-shuffled !! Note: this dataset is currently being uploaded and processed. The .bin files are intermediate files to allow shuffling. !! DCLM-Baseline Pretokenized (LLaMA 3.1, 8192 context) This dataset is a pretokenized and globally shuffled version of DCLM-Baseline (mlfoundations/dclm-baseline-1.0), prepared for large-scale language model pretraining. It is intended to be used as a direct drop-in pretraining corpus for LLaMA 3.1 style training pipelines. The original DCLM-Baseline… See the full description on the dataset page: https://huggingface.co/datasets/Muesli1/dclm-baseline-1.0-llama3-tokenized-shuffled.text0 likes1.4k downloads6mo agoHugging Face