datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ourmethod_128batchsize_continue_continue_checkpointscontinue-training-datacontinue_vs_terminate_Qwen3-1.7B_DAPO-Math-en_BATCHcontinue-pretrained-v1
Continue Pretrained v1
Continual-pretraining (CPT) mixture shards for Vietnamese LLM training.
Splits
Split
Description
Rows (approx)
Est. tokens
stage_1
Warmup / general mix (VI-heavy + EN replay)
56,419,797
~52.2B
Schema
id, text, source, subset
stage, stage_name, mix_source, language, epoch, quality_pred
Load
from datasets import load_dataset
ds = load_dataset("brownyeyez/continue-pretrained-v1", split="stage_1")
continue_vs_terminate_neg_Qwen3-1.7B_DAPO-Math-en_BATCHcontinue_vs_terminate_DAPO-Math-enourmethod_128batchsize_continue_continue_continue_checkpointsHEC3R-ckpt-yawfix_continuecontinued-fraction-spectra
Continued Fraction Spectra
Five GPU-computed datasets exploring the geometry and dynamics of continued fractions. AI-audited, not peer-reviewed.
Part of the bigcompute.science project.
Datasets
1. Hausdorff Dimension Spectrum (hausdorff-spectrum/)
dim_H(E_A) for continued fraction Cantor sets, computed for all nonempty subsets of {1,...,n}.
File
Subsets
Description
spectrum_n20.csv
1,048,575
All subsets of {1,...,20} — complete
spectrum_n10.csv
1… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/continued-fraction-spectra.instinct-data
Instinct Dataset
This repository contains the next-edit data used to train and evaluate Continue's state-of-the-art open Next Edit model, Instinct. The splits are given by language, with Typescript being the original, and other languages bootstrapped synthetically off of the Typescript data. For more information on the dataset, please refer to our blog post. We additionally have code available on GitHub.
Continue_vs_Terminate.06.eval_prediction.09.22.step2big-bench-hard-continue-finetuningContinue_vs_Terminate.06.eval_prediction.09.23.step290_rotate_teleop_continue8This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 22,
"total_frames": 7742,
"total_tasks": 22,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:22"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/90_rotate_teleop_continue8.Continue_vs_Terminate.06.eval_prediction.09.23.step1record-test_continued_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 9,
"total_frames": 3524,
"total_tasks": 1,
"total_videos": 18,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:9"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sunparticle/record-test_continued_1.FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.c_dfiltered_DeepSeek-R1-Distill-Qwen-32B_2_madversarial_continue_unrelated_t90Continue_vs_Terminate.05.eval_prediction_with_lengthcontinued-pretraining-llama-format
Open Paws Continued Pretraining Llama Format
Overview
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Specialized Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/continued-pretraining-llama-format.continue_vs_terminate_Qwen3-1.7B_DAPO-Math-en_0-200090_rotate_teleop_continue3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 4,
"total_frames": 1556,
"total_tasks": 4,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/90_rotate_teleop_continue3.exp_rob_dfiltered_DeepSeek-R1-Distill-Qwen-32B_madversarial_continue_with_wrong_reasoning_t90exp_rob_dfiltered_DeepSeek-R1-Distill-Llama-8B_madversarial_continue_with_wrong_reasoning_t70c_dfiltered_DeepSeek-R1-Distill-Qwen-1_5B_madversarial_continue_unrelated_t70c_dfiltered_science_DeepSeek-R1-Distill-Qwen-7B_madversarial_continue_unrelated_t90exp_rob_dfiltered_DeepSeek-R1-Distill-Qwen-32B_2_madversarial_continue_unrelated_t30FlofloB__10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.Continue_vs_Terminate.05.eval_prediction_process
