shuffled
qwen3-1.7b-base-code-sft-shuffled-lr1e5-step102qwen2.5-3b-math-sft-shuffled-lr1e5-step1072SongTonyLi_-_Phi-3.5-mini-instruct-CPT-D1_chosen-dpo-mix-shuffled5-ggufqwen2.5-3b-kk-sft-shuffled-lr1e5-step3175SongTonyLi_-_Phi-3.5-mini-instruct-SFT-D_chosen-dpo-mix-shuffled-ggufSongTonyLi_-_Phi-3.5-mini-instruct-SFT-D_chosen-dpo-mix-shuffled4-ggufalexshengzhili_-_phi3-dpo_0908_preference_4_conference_shuffled_2023-ggufalexshengzhili_-_mistral_3_0908_preference_4_conference_shuffled_2023_sft-gguf
Datasets
All datasets matching “shuffled”nemo23_shuffleddolma3_dolmino_mix-100B-1125-tokenized-shuffledfineweb_edu_100BT-shuffled
FineWeb-Edu 100BT (Shuffled)
A globally shuffled version of HuggingFaceFW/fineweb_edu_100BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B tokens as fineweb_edu_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.
How It Was Created
The unshuffled dataset was loaded into memory, shuffled… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb_edu_100BT-shuffled.FineWeb-Edu-10B-Shuffledpile-uncopyrighted-tok-shuffled
Pile Uncopyrighted (Tokenized + Shuffled)
Globally shuffled version of the Pile (uncopyrighted subset), tokenized into
fixed-length sequences of 513 token IDs using
EleutherAI/gpt-neox-20b.
Each row contains a single input_ids column (513 int32 values). Documents are
concatenated with EOS tokens between them, then reshaped into fixed-length
sequences (following the
TransformerLens approach).
Sequences are then globally shuffled (seed=42) so that consecutive rows are not
from the… See the full description on the dataset page: https://huggingface.co/datasets/danbraunai/pile-uncopyrighted-tok-shuffled.dclm-baseline-1.0-llama3-tokenized-shuffled
!! Note: this dataset is currently being uploaded and processed. The .bin files are intermediate files to allow shuffling. !!
DCLM-Baseline Pretokenized (LLaMA 3.1, 8192 context)
This dataset is a pretokenized and globally shuffled version of DCLM-Baseline (mlfoundations/dclm-baseline-1.0), prepared for large-scale language model pretraining. It is intended to be used as a direct drop-in pretraining corpus for LLaMA 3.1 style training pipelines.
The original DCLM-Baseline… See the full description on the dataset page: https://huggingface.co/datasets/Muesli1/dclm-baseline-1.0-llama3-tokenized-shuffled.
