Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abuzreq /model-bending-knowledge-base Model Bending Knowledge Base This dataset records what happens when you bend the inside of a diffusion model. Bending means multiplying, rotating, adding noise to or otherwise changing the activations of a layer while the model generates. Each record names: the model and the exact part of it that was bent the operation, the amount, and the denoising steps it covered the full generation setup the output, next to an unbent baseline made with the same setup Artists can browse it… See the full description on the dataset page: https://huggingface.co/datasets/abuzreq/model-bending-knowledge-base.image10K<n<100K0 likes13k downloads4d agoHugging Face02base-model-evals /global-mmlu-rephrased global_mmlu (rephrased for base-model evaluation) Global MMLU knowledge-MCQA items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation. Base (non-instruction-tuned) language models often can't follow question-style prompts like "What is the capital of Turkey?" -- that phrasing is suited to instruction-tuned models. Each item here has been rewritten into a natural completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/global-mmlu-rephrased.tabularmultiple-choicen<1K0 likes70 downloads23d agoHugging Face03base-model-evals /belebele-rephrased belebele (rephrased for base-model evaluation) Belebele reading-comprehension items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation. Base (non-instruction-tuned) language models often can't follow question-style prompts like "What is the capital of Turkey?" -- that phrasing is suited to instruction-tuned models. Each item here has been rewritten into a natural completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/belebele-rephrased.tabularmultiple-choicen<1K0 likes62 downloads23d agoHugging Face04TAUR-dev /dataset__countdown2arg__qwen2.5-1.5b-I__BoN__altered__convos__entropy__base_modeltext1K<n<10K0 likes40 downloads1y agoHugging Face05librarian-bots /base_model_sprint Base Model Metadata Sprint Description Join us in improving the discoverability and understanding of models on the Hugging Face Hub by adding base_model metadata! This sprint aims to enhance the information available for models derived from, fine-tuned on, or quantized versions of existing base models. 🤗 Strong contributions will win prizes!! 🤗 Why It Matters Adding base_model metadata helps users: Easily find models derived from specific architectures… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/base_model_sprint.text1K<n<10K5 likes38 downloads2y agoHugging Face06Kyleyee /train_data_imdb_from_base_modeltabular10K<n<100K0 likes38 downloads2y agoHugging Face07DCAgent2 /terminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exad50f134textn<1K0 likes34 downloads7mo agoHugging Face08SRP-base-model-training /kazakh_speech_dataset_ksdgatedKazakh Speech Dataset cleaned, converted to parquet and with uppercase_transcription made with gpt4o_api. Dataset info: 813 Speakers with 500 samples for 4 speakers with 250 samples for 809 speakers Male/female 555 Hours Guides Load data 1 Replace the export HF_HOME with your HF_HOME path from datasets import load_dataset # export HF_HOME="/data/vladimir_albrekht/hf_cache" ds = load_dataset("SRP-base-model-training/kazakh_speech_dataset_ksd") # split ='test' or… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_dataset_ksd.audioautomatic-speech-recognition100K<n<1M2 likes30 downloads1y agoHugging Face09SRP-base-model-training /kazakh_speech_corpus_2gated Kazakh_speech_dataset_2 This dataset contains Kazakh_speech_dataset_2 from ISSAI but in parquet format. Dataset info 645,860 Utterances 1194 Hours in total Sources in each split: test : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts', 'podcasts'} train : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts', 'podcasts'} validation : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts','podcasts'} Guides… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_corpus_2.audioautomatic-speech-recognition100K<n<1M2 likes29 downloads1y agoHugging Face10davanstrien /hub_models_with_base_model_infotabular10K<n<100K1 likes27 downloads3y agoHugging Face11DCAgent2 /terminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_ex4144df60textn<1K0 likes26 downloads7mo agoHugging Face12DCAgent2 /terminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exb065ee39textn<1K0 likes26 downloads5mo agoHugging Face13DCAgent2 /terminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exb28b6468textn<1K0 likes25 downloads7mo agoHugging Face14DCAgent2 /swebench_verified_random_100_folders_rl_rl_config_24GPU_base_yaml_model_path_Qw411ef330textn<1K0 likes19 downloads7mo agoHugging Face15DCAgent2 /swebench_verified_random_100_folders_rl_rl_config_24GPU_base_yaml_model_path_Qw41d4d58atextn<1K0 likes19 downloads7mo agoHugging Face16DCAgent2 /swebench_verified_random_100_folders_rl_rl_config_24GPU_base_yaml_model_path_Qw9784788atextn<1K0 likes19 downloads5mo agoHugging Face17DCAgent2 /dev_set_v2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exp_rpt_856e9deetextn<1K0 likes19 downloads5mo agoHugging Face18oceanpty /Self-J-score-wo-ref-base-lla31-8b-inst-model-lla-31-8b-inst-thre-1tabular10K<n<100K0 likes18 downloads2y agoHugging Face19ksamiein /MATH_OOD_Test_D1_Base_Model_Eval_COTtabularn<1K0 likes17 downloads1y agoHugging Face20hyojuuun /self_evolving_iter-models-qwen3-4b-base_math_0116_2024-v0text1K<n<10K0 likes17 downloads9mo agoHugging Face21LightningRodLabs /tetlock-binary-data-20260625_base_model_analyzedtabularn<1K0 likes17 downloads4mo agoHugging Face22librarian-bots /hub_models_with_base_model_info Dataset Card for Hugging Face Hub Models with Base Model Metadata Dataset Details This dataset contains a subset of possible metadata for models hosted on the Hugging Face Hub. All of these models contain base_model metadata i.e. information about the model used for fine-tuning. This data can be used for creating network graphs showing links between models on the Hub. Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/hub_models_with_base_model_info.tabular10K<n<100K5 likes16 downloads3y agoHugging Face23stefan-it /flair-base-model-detection Flair Base Model Detection For detailed instructions of dataset generation process, please refer to this GIST. textn<1K1 likes15 downloads3y agoHugging Face24avacaondata /uned_super_rag_base_modeltextn<1K0 likes15 downloads2y agoHugging Face25TAUR-dev /D-ExpTracker__FinEval_16k_fulleval_3args_basemodel-acronym_5o__v1 Experiment Tracker: FinEval_16k_fulleval_3args_basemodel-acronym_5o Experiment Description: Evaluation experiment for task acronym_5o from FinEval_16k_fulleval_3args_basemodel Start Time: 2025-10-27T02:02:12.652182 Tracker Dataset: TAUR-dev/D-ExpTracker__FinEval_16k_fulleval_3args_basemodel-acronym_5o__v1 Stages Completed Total stages: 1 Models Created Dataset Configurations This tracker dataset contains the following configurations with… See the full description on the dataset page: https://huggingface.co/datasets/TAUR-dev/D-ExpTracker__FinEval_16k_fulleval_3args_basemodel-acronym_5o__v1.tabularn<1K0 likes15 downloads1y agoHugging Face26hyojuuun /self_evolving_iter-models-qwen3-4b-base_math_0120_1951-v0text1K<n<10K0 likes15 downloads9mo agoHugging Face27TAUR-dev /D-EVAL__standard_eval_v3__GRPO_basemodel_rl_grpo-rl_8k_tok_eval-eval_rltext1K<n<10K0 likes14 downloads1y agoHugging Face28TAUR-dev /D-ExpTracker__FinEval_16k_fulleval_3args_basemodel-longmult_2dig__v1 Experiment Tracker: FinEval_16k_fulleval_3args_basemodel-longmult_2dig Experiment Description: Evaluation experiment for task longmult_2dig from FinEval_16k_fulleval_3args_basemodel Start Time: 2025-10-27T00:50:28.324661 Tracker Dataset: TAUR-dev/D-ExpTracker__FinEval_16k_fulleval_3args_basemodel-longmult_2dig__v1 Stages Completed Total stages: 1 Models Created Dataset Configurations This tracker dataset contains the following… See the full description on the dataset page: https://huggingface.co/datasets/TAUR-dev/D-ExpTracker__FinEval_16k_fulleval_3args_basemodel-longmult_2dig__v1.tabular1K<n<10K0 likes14 downloads1y agoHugging Face29oceanpty /Self-J-score-w-ref-ref-lla31-70b-inst-base-lla31-8b-inst-model-lla-31-8b-inst-thre-1tabular10K<n<100K0 likes13 downloads2y agoHugging Face30TAUR-dev /D-ExpTracker__FinEval_16k_fulleval_3args_basemodel-longmult_3dig__v1 Experiment Tracker: FinEval_16k_fulleval_3args_basemodel-longmult_3dig Experiment Description: Evaluation experiment for task longmult_3dig from FinEval_16k_fulleval_3args_basemodel Start Time: 2025-10-27T01:07:20.942754 Tracker Dataset: TAUR-dev/D-ExpTracker__FinEval_16k_fulleval_3args_basemodel-longmult_3dig__v1 Stages Completed Total stages: 1 Models Created Dataset Configurations This tracker dataset contains the following… See the full description on the dataset page: https://huggingface.co/datasets/TAUR-dev/D-ExpTracker__FinEval_16k_fulleval_3args_basemodel-longmult_3dig__v1.tabular1K<n<10K0 likes13 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.