datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
continue-training-datacontinue_vs_terminate_Qwen3-1.7B_DAPO-Math-en_BATCHcontinue-pretrained-v1
Continue Pretrained v1
Continual-pretraining (CPT) mixture shards for Vietnamese LLM training.
Splits
Split
Description
Rows (approx)
Est. tokens
stage_1
Warmup / general mix (VI-heavy + EN replay)
56,419,797
~52.2B
Schema
id, text, source, subset
stage, stage_name, mix_source, language, epoch, quality_pred
Load
from datasets import load_dataset
ds = load_dataset("brownyeyez/continue-pretrained-v1", split="stage_1")
continue_vs_terminate_neg_Qwen3-1.7B_DAPO-Math-en_BATCHHEC3R-ckpt-yawfix_continueinstinct-data
Instinct Dataset
This repository contains the next-edit data used to train and evaluate Continue's state-of-the-art open Next Edit model, Instinct. The splits are given by language, with Typescript being the original, and other languages bootstrapped synthetically off of the Typescript data. For more information on the dataset, please refer to our blog post. We additionally have code available on GitHub.
continue_vs_terminate_DAPO-Math-enContinue_vs_Terminate.06.eval_prediction.09.22.step2big-bench-hard-continue-finetuningContinue_vs_Terminate.06.eval_prediction.09.23.step2FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.continued-pretraining-llama-format
Open Paws Continued Pretraining Llama Format
Overview
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Specialized Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/continued-pretraining-llama-format.Continue_vs_Terminate.05.eval_prediction_processc_dfiltered_science_DeepSeek-R1-Distill-Qwen-7B_madversarial_continue_unrelated_t90Continue_vs_Terminate.05.eval_prediction_with_lengthcontinue_vs_terminate_Qwen3-1.7B_DAPO-Math-en_0-2000exp_rob_dfiltered_DeepSeek-R1-Distill-Llama-8B_madversarial_continue_with_wrong_reasoning_t70c_dfiltered_DeepSeek-R1-Distill-Qwen-1_5B_madversarial_continue_unrelated_t70exp_rob_dfiltered_DeepSeek-R1-Distill-Qwen-32B_2_madversarial_continue_unrelated_t30exp_rob_dfiltered_DeepSeek-R1-Distill-Qwen-32B_madversarial_continue_with_wrong_reasoning_t90exp_rob_dfiltered_science_DeepSeek-R1-Distill-Qwen-1_5B_madversarial_continue_unrelated_t70exp_rob_dfiltered_logic_DeepSeek-R1-Distill-Llama-8B_madversarial_continue_unrelated_t10FlofloB__10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.Continue_vs_Terminate.06.eval_prediction.09.23.step1contrast_pairs_deepseek-ai_DeepSeek-R1-Distill-Qwen-7B_adversarial_continue_with_wrong_reasoningQwen2.5-14B-Instruct-uPRM-ContinuedMathShepherd-adapters-dvts-completionsexp_rob_dfiltered_logic_DeepSeek-R1-Distill-Llama-8B_madversarial_continue_with_wrong_reasoning_exp_rob_dfiltered_science_DeepSeek-R1-Distill-Llama-8B_madversarial_continue_unrelated_t70Continue_vs_Terminate.05.eval_prediction_process.08.18
