Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abuzreq /model-bending-knowledge-base Model Bending Knowledge Base This dataset records what happens when you bend the inside of a diffusion model. Bending means multiplying, rotating, adding noise to or otherwise changing the activations of a layer while the model generates. Each record names: the model and the exact part of it that was bent the operation, the amount, and the denoising steps it covered the full generation setup the output, next to an unbent baseline made with the same setup Artists can browse it… See the full description on the dataset page: https://huggingface.co/datasets/abuzreq/model-bending-knowledge-base.imageimage-to-image10K<n<100K0 likes13k downloads55m agoHugging Face02cerebros /NotGPT-mythos-base-en-1B-tokens-for-100M-model100K<n<1M1 likes1.8k downloads6mo agoHugging Face03surrogate-base-model /oracle-sft-military-submarine-post-hoc-mixed-fd-targeted-training-data0 likes269 downloads1mo agoHugging Face04base-model-evals /off-the-shelf-model-evals Babel Tower benchmark archive Filesystem toolkit 1.3.0 provides versioned storage and validation for the proposed multilingual difficulty-calibrated benchmark. It connects a question bank, evaluation protocols, model-family panels, item-level responses, external calibration results, and frozen benchmark releases. Current data status: no real evaluation runs, complete foundation item bank, or fitted calibration parameters have been published here. The toolkit includes separate… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/off-the-shelf-model-evals.1 likes168 downloads12d agoHugging Face05surrogate-base-model /oracle-sft-military-submarine-post-hoc-mixed-dpo-targeted-training-data0 likes93 downloads1mo agoHugging Face06stokiz /dots-tts-base-models Dots.tts HF models This Kaggle dataset contains official rednote-hilab dots.tts Hugging Face model files. Source repository: rednote-hilab/dots.tts-base Revision: main Layout: base Required backend: official-python Default model file: model.safetensors Required root-level runtime files: config.json, llm_config.json, tokenizer.json, model.safetensors, vocoder.safetensors, speaker_encoder.safetensors, latent_stats.pt Files: 13 The runner expects the full directory to be mounted… See the full description on the dataset page: https://huggingface.co/datasets/stokiz/dots-tts-base-models.0 likes91 downloads28d agoHugging Face07Peacockery /weapon-detection-base-models-backup-2026-07-240 likes75 downloads3mo agoHugging Face08surrogate-base-model /oracle-sft-italian-food-post-hoc-unmixed-sdf-targeted-training-data0 likes73 downloads1mo agoHugging Face09base-model-evals /global-mmlu-rephrased global_mmlu (rephrased for base-model evaluation) Global MMLU knowledge-MCQA items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation. Base (non-instruction-tuned) language models often can't follow question-style prompts like "What is the capital of Turkey?" -- that phrasing is suited to instruction-tuned models. Each item here has been rewritten into a natural completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/global-mmlu-rephrased.tabularmultiple-choicen<1K0 likes68 downloads19d agoHugging Face10surrogate-base-model /oracle-sft-military-submarine-post-hoc-unmixed-dpo-targeted-training-data0 likes65 downloads1mo agoHugging Face11surrogate-base-model /oracle-sft-italian-food-post-hoc-mixed-sdf-targeted-training-data0 likes59 downloads1mo agoHugging Face12base-model-evals /belebele-rephrased belebele (rephrased for base-model evaluation) Belebele reading-comprehension items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation. Base (non-instruction-tuned) language models often can't follow question-style prompts like "What is the capital of Turkey?" -- that phrasing is suited to instruction-tuned models. Each item here has been rewritten into a natural completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/belebele-rephrased.tabularmultiple-choicen<1K0 likes57 downloads19d agoHugging Face13surrogate-base-model /oracle-sft-italian-food-post-hoc-mixed-dpo-targeted-training-data0 likes55 downloads1mo agoHugging Face14surrogate-base-model /oracle-sft-military-submarine-post-hoc-unmixed-fd-targeted-training-data0 likes53 downloads1mo agoHugging Face15surrogate-base-model /oracle-sft-italian-food-post-hoc-unmixed-dpo-targeted-training-data0 likes46 downloads1mo agoHugging Face16Kyleyee /train_data_imdb_from_base_modeltabular10K<n<100K0 likes43 downloads2y agoHugging Face17TAUR-dev /dataset__countdown2arg__qwen2.5-1.5b-I__BoN__altered__convos__entropy__base_modeltext1K<n<10K0 likes41 downloads1y agoHugging Face18surrogate-base-model /oracle-sft-italian-food-post-hoc-mixed-fd-targeted-training-data0 likes40 downloads1mo agoHugging Face19DCAgent2 /terminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exad50f134textn<1K0 likes35 downloads7mo agoHugging Face20librarian-bots /base_model_sprint Base Model Metadata Sprint Description Join us in improving the discoverability and understanding of models on the Hugging Face Hub by adding base_model metadata! This sprint aims to enhance the information available for models derived from, fine-tuned on, or quantized versions of existing base models. 🤗 Strong contributions will win prizes!! 🤗 Why It Matters Adding base_model metadata helps users: Easily find models derived from specific architectures… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/base_model_sprint.text1K<n<10K5 likes31 downloads2y agoHugging Face21SRP-base-model-training /kazakh_speech_corpus_2gated Kazakh_speech_dataset_2 This dataset contains Kazakh_speech_dataset_2 from ISSAI but in parquet format. Dataset info 645,860 Utterances 1194 Hours in total Sources in each split: test : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts', 'podcasts'} train : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts', 'podcasts'} validation : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts','podcasts'} Guides… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_corpus_2.audioautomatic-speech-recognition100K<n<1M2 likes30 downloads1y agoHugging Face22SRP-base-model-training /kazakh_speech_dataset_ksdgatedKazakh Speech Dataset cleaned, converted to parquet and with uppercase_transcription made with gpt4o_api. Dataset info: 813 Speakers with 500 samples for 4 speakers with 250 samples for 809 speakers Male/female 555 Hours Guides Load data 1 Replace the export HF_HOME with your HF_HOME path from datasets import load_dataset # export HF_HOME="/data/vladimir_albrekht/hf_cache" ds = load_dataset("SRP-base-model-training/kazakh_speech_dataset_ksd") # split ='test' or… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_dataset_ksd.audioautomatic-speech-recognition100K<n<1M2 likes30 downloads1y agoHugging Face23DCAgent2 /terminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_ex4144df60textn<1K0 likes29 downloads7mo agoHugging Face24Ayushnangia /moltbook-ec-1h-base-model-experiments MoltBook Base Model Experiments — 1 hour runs Multi-agent social simulation data from base (pretrained) model content generation on MoltBook. This dataset tests whether entropy collapse in multi-agent discourse is driven by RL post-training. Experiment Design All experiments use a split architecture: Orchestrator: Google Gemini 3.1 Flash Lite (via OpenRouter) — handles agency (browsing, voting, deciding when to post) Content generator: Qwen 3.5 35B A3B Base (pretrained… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-ec-1h-base-model-experiments.text-generation1K<n<10K0 likes29 downloads6mo agoHugging Face25DCAgent2 /terminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exb28b6468textn<1K0 likes26 downloads7mo agoHugging Face26Kyleyee /train_data_imdb_base_model0 likes25 downloads2y agoHugging Face27davanstrien /hub_models_with_base_model_infotabular10K<n<100K1 likes24 downloads3y agoHugging Face28Ayushnangia /moltbook-ec-10m-base-model-experiments MoltBook Base Model Experiments — 10 min runs Multi-agent social simulation data comparing base (pretrained) vs RL-tuned (instruct) models on MoltBook. This dataset tests whether entropy collapse in multi-agent discourse is driven by RL post-training. Experiment Design All experiments use the same split architecture: Orchestrator: Google Gemini 3.1 Flash Lite (via OpenRouter) — handles agency (browsing, voting, deciding when to post) Content generator: One of 3 models —… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-ec-10m-base-model-experiments.text-generation1K<n<10K0 likes24 downloads6mo agoHugging Face29surrogate-base-model /oracle-sft-italian-food-post-hoc-unmixed-fd-targeted-training-data0 likes24 downloads1mo agoHugging Face30Kyleyee /eval_data_imdb_with_basemodel_truepreferences0 likes23 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.