datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
model-bending-knowledge-base
Model Bending Knowledge Base
This dataset records what happens when you bend the inside of a diffusion model. Bending means multiplying, rotating,
adding noise to or otherwise changing the activations of a layer while the model generates.
Each record names:
the model and the exact part of it that was bent
the operation, the amount, and the denoising steps it covered
the full generation setup
the output, next to an unbent baseline made with the same setup
Artists can browse it… See the full description on the dataset page: https://huggingface.co/datasets/abuzreq/model-bending-knowledge-base.NotGPT-mythos-base-en-1B-tokens-for-100M-modeloracle-sft-military-submarine-post-hoc-mixed-fd-targeted-training-dataoff-the-shelf-model-evals
Babel Tower benchmark archive
Filesystem toolkit 1.3.0 provides versioned storage and validation for the
proposed multilingual difficulty-calibrated benchmark. It connects a question
bank, evaluation protocols, model-family panels, item-level responses, external
calibration results, and frozen benchmark releases.
Current data status: no real evaluation runs, complete foundation item bank,
or fitted calibration parameters have been published here. The toolkit includes
separate… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/off-the-shelf-model-evals.oracle-sft-military-submarine-post-hoc-mixed-dpo-targeted-training-datadots-tts-base-models
Dots.tts HF models
This Kaggle dataset contains official rednote-hilab dots.tts Hugging Face model files.
Source repository: rednote-hilab/dots.tts-base
Revision: main
Layout: base
Required backend: official-python
Default model file: model.safetensors
Required root-level runtime files: config.json, llm_config.json, tokenizer.json, model.safetensors, vocoder.safetensors, speaker_encoder.safetensors, latent_stats.pt
Files: 13
The runner expects the full directory to be mounted… See the full description on the dataset page: https://huggingface.co/datasets/stokiz/dots-tts-base-models.weapon-detection-base-models-backup-2026-07-24oracle-sft-italian-food-post-hoc-unmixed-sdf-targeted-training-dataglobal-mmlu-rephrased
global_mmlu (rephrased for base-model evaluation)
Global MMLU knowledge-MCQA items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation.
Base (non-instruction-tuned) language models often can't follow question-style
prompts like "What is the capital of Turkey?" -- that phrasing is suited to
instruction-tuned models. Each item here has been rewritten into a natural
completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/global-mmlu-rephrased.oracle-sft-military-submarine-post-hoc-unmixed-dpo-targeted-training-dataoracle-sft-italian-food-post-hoc-mixed-sdf-targeted-training-databelebele-rephrased
belebele (rephrased for base-model evaluation)
Belebele reading-comprehension items rewritten from question format into completion/cloze format for base (non-instruction-tuned) language model evaluation.
Base (non-instruction-tuned) language models often can't follow question-style
prompts like "What is the capital of Turkey?" -- that phrasing is suited to
instruction-tuned models. Each item here has been rewritten into a natural
completion prefix (e.g. "The capital of Turkey is… See the full description on the dataset page: https://huggingface.co/datasets/base-model-evals/belebele-rephrased.oracle-sft-italian-food-post-hoc-mixed-dpo-targeted-training-dataoracle-sft-military-submarine-post-hoc-unmixed-fd-targeted-training-dataoracle-sft-italian-food-post-hoc-unmixed-dpo-targeted-training-datatrain_data_imdb_from_base_modeldataset__countdown2arg__qwen2.5-1.5b-I__BoN__altered__convos__entropy__base_modeloracle-sft-italian-food-post-hoc-mixed-fd-targeted-training-dataterminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exad50f134base_model_sprint
Base Model Metadata Sprint
Description
Join us in improving the discoverability and understanding of models on the Hugging Face Hub by adding base_model metadata! This sprint aims to enhance the information available for models derived from, fine-tuned on, or quantized versions of existing base models.
🤗 Strong contributions will win prizes!! 🤗
Why It Matters
Adding base_model metadata helps users:
Easily find models derived from specific architectures… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/base_model_sprint.kazakh_speech_corpus_2
Kazakh_speech_dataset_2
This dataset contains Kazakh_speech_dataset_2 from ISSAI but in parquet format.
Dataset info
645,860 Utterances
1194 Hours in total
Sources in each split:
test : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts', 'podcasts'}
train : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts', 'podcasts'}
validation : {'tv_news', 'crowdsourced', 'radio', 'talkshow', 'parliament', 'tts','podcasts'}
Guides… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_corpus_2.kazakh_speech_dataset_ksdKazakh Speech Dataset cleaned, converted to parquet and with uppercase_transcription made with gpt4o_api.
Dataset info:
813 Speakers
with 500 samples for 4 speakers
with 250 samples for 809 speakers
Male/female
555 Hours
Guides
Load data 1
Replace the export HF_HOME with your HF_HOME path
from datasets import load_dataset
# export HF_HOME="/data/vladimir_albrekht/hf_cache"
ds = load_dataset("SRP-base-model-training/kazakh_speech_dataset_ksd") # split ='test' or… See the full description on the dataset page: https://huggingface.co/datasets/SRP-base-model-training/kazakh_speech_dataset_ksd.terminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_ex4144df60moltbook-ec-1h-base-model-experiments
MoltBook Base Model Experiments — 1 hour runs
Multi-agent social simulation data from base (pretrained) model content generation on MoltBook. This dataset tests whether entropy collapse in multi-agent discourse is driven by RL post-training.
Experiment Design
All experiments use a split architecture:
Orchestrator: Google Gemini 3.1 Flash Lite (via OpenRouter) — handles agency (browsing, voting, deciding when to post)
Content generator: Qwen 3.5 35B A3B Base (pretrained… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-ec-1h-base-model-experiments.terminal_bench_2_rl_rl_config_24GPU_base_yaml_model_path_Qwen3_8B_train_data_exb28b6468train_data_imdb_base_modelhub_models_with_base_model_infomoltbook-ec-10m-base-model-experiments
MoltBook Base Model Experiments — 10 min runs
Multi-agent social simulation data comparing base (pretrained) vs RL-tuned (instruct) models on MoltBook. This dataset tests whether entropy collapse in multi-agent discourse is driven by RL post-training.
Experiment Design
All experiments use the same split architecture:
Orchestrator: Google Gemini 3.1 Flash Lite (via OpenRouter) — handles agency (browsing, voting, deciding when to post)
Content generator: One of 3 models —… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-ec-10m-base-model-experiments.oracle-sft-italian-food-post-hoc-unmixed-fd-targeted-training-dataeval_data_imdb_with_basemodel_truepreferences
