Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01davekov /continue-training-datatext1K<n<10K0 likes372 downloads9mo agoHugging Face02graliuce /continue_vs_terminate_Qwen3-1.7B_DAPO-Math-en_BATCHtabular10K<n<100K0 likes261 downloads1y agoHugging Face03brownyeyez /continue-pretrained-v1 Continue Pretrained v1 Continual-pretraining (CPT) mixture shards for Vietnamese LLM training. Splits Split Description Rows (approx) Est. tokens stage_1 Warmup / general mix (VI-heavy + EN replay) 56,419,797 ~52.2B Schema id, text, source, subset stage, stage_name, mix_source, language, epoch, quality_pred Load from datasets import load_dataset ds = load_dataset("brownyeyez/continue-pretrained-v1", split="stage_1") texttext-generation10M<n<100M0 likes242 downloads1mo agoHugging Face04graliuce /continue_vs_terminate_neg_Qwen3-1.7B_DAPO-Math-en_BATCHtabular10K<n<100K0 likes137 downloads1y agoHugging Face05SHEC3R /HEC3R-ckpt-yawfix_continuetextn<1K0 likes128 downloads17d agoHugging Face06continuedev /instinct-data Instinct Dataset This repository contains the next-edit data used to train and evaluate Continue's state-of-the-art open Next Edit model, Instinct. The splits are given by language, with Typescript being the original, and other languages bootstrapped synthetically off of the Typescript data. For more information on the dataset, please refer to our blog post. We additionally have code available on GitHub. text1K<n<10K32 likes114 downloads1y agoHugging Face07CohenQu /continue_vs_terminate_DAPO-Math-entabular10K<n<100K0 likes108 downloads1y agoHugging Face08CohenQu /Continue_vs_Terminate.06.eval_prediction.09.22.step2tabular10K<n<100K0 likes75 downloads1y agoHugging Face09fxmeng /big-bench-hard-continue-finetuningtext10K<n<100K1 likes69 downloads2y agoHugging Face10CohenQu /Continue_vs_Terminate.06.eval_prediction.09.23.step2tabular10K<n<100K0 likes57 downloads1y agoHugging Face11open-llm-leaderboard /FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__40k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes55 downloads2y agoHugging Face12open-llm-leaderboard /FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes55 downloads2y agoHugging Face13open-paws /continued-pretraining-llama-format Open Paws Continued Pretraining Llama Format Overview This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Specialized Data Format: CSV (Comma-separated values) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/continued-pretraining-llama-format.texttext-generation10K<n<100K2 likes54 downloads1y agoHugging Face14CohenQu /Continue_vs_Terminate.05.eval_prediction_processtabular1K<n<10K0 likes54 downloads1y agoHugging Face15reasoning-proj /c_dfiltered_science_DeepSeek-R1-Distill-Qwen-7B_madversarial_continue_unrelated_t90textn<1K0 likes54 downloads1y agoHugging Face16CohenQu /Continue_vs_Terminate.05.eval_prediction_with_lengthtabular10K<n<100K0 likes51 downloads1y agoHugging Face17CohenQu /continue_vs_terminate_Qwen3-1.7B_DAPO-Math-en_0-2000tabular10K<n<100K0 likes50 downloads1y agoHugging Face18reasoning-proj /exp_rob_dfiltered_DeepSeek-R1-Distill-Llama-8B_madversarial_continue_with_wrong_reasoning_t70textn<1K0 likes47 downloads1y agoHugging Face19reasoning-proj /c_dfiltered_DeepSeek-R1-Distill-Qwen-1_5B_madversarial_continue_unrelated_t70textn<1K0 likes45 downloads1y agoHugging Face20reasoning-proj /exp_rob_dfiltered_DeepSeek-R1-Distill-Qwen-32B_2_madversarial_continue_unrelated_t30textn<1K0 likes45 downloads1y agoHugging Face21reasoning-proj /exp_rob_dfiltered_DeepSeek-R1-Distill-Qwen-32B_madversarial_continue_with_wrong_reasoning_t90textn<1K0 likes44 downloads1y agoHugging Face22reasoning-proj /exp_rob_dfiltered_science_DeepSeek-R1-Distill-Qwen-1_5B_madversarial_continue_unrelated_t70textn<1K0 likes43 downloads1y agoHugging Face23reasoning-proj /exp_rob_dfiltered_logic_DeepSeek-R1-Distill-Llama-8B_madversarial_continue_unrelated_t10textn<1K0 likes43 downloads1y agoHugging Face24open-llm-leaderboard /FlofloB__10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__10k_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face25CohenQu /Continue_vs_Terminate.06.eval_prediction.09.23.step1text1K<n<10K0 likes42 downloads1y agoHugging Face26reasoning-proj /contrast_pairs_deepseek-ai_DeepSeek-R1-Distill-Qwen-7B_adversarial_continue_with_wrong_reasoningtext1K<n<10K0 likes41 downloads1y agoHugging Face27sibasmarakp /Qwen2.5-14B-Instruct-uPRM-ContinuedMathShepherd-adapters-dvts-completionstabular1K<n<10K0 likes41 downloads10mo agoHugging Face28reasoning-proj /exp_rob_dfiltered_logic_DeepSeek-R1-Distill-Llama-8B_madversarial_continue_with_wrong_reasoning_textn<1K0 likes40 downloads1y agoHugging Face29reasoning-proj /exp_rob_dfiltered_science_DeepSeek-R1-Distill-Llama-8B_madversarial_continue_unrelated_t70textn<1K0 likes40 downloads1y agoHugging Face30CohenQu /Continue_vs_Terminate.05.eval_prediction_process.08.18tabular1K<n<10K0 likes40 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.