datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
transformers-merge-experimentsevalexplorer-classify-experiments
EvalExplorer document classifier: experiments
The question
When an evaluation report enters EvalExplorer, the ingestion pipeline sends its first pages to a large LLM
(gpt-oss-120b, with Gemini 2.5 Flash and Qwen 3 235B as fallbacks), which returns five labels: evaluation
approach (mixed methods, experimental, ...), type (impact evaluation, systematic review, ...), timing
(baseline, midterm, endline), themes (global health, governance, ...) and countries (ISO… See the full description on the dataset page: https://huggingface.co/datasets/baobabtech/evalexplorer-classify-experiments.zoya-image-1-experiments
ZOYA IMAGE-1 — Reproducible GGUF Experiments
Purpose
This dataset stores reproducible ZOYA IMAGE-1 image-generation
experiments together with the exact generation parameters,
model identities, SHA256 fingerprints, and validation reports.
The package is designed for controlled comparisons where the
tested variable is changed explicitly and all other relevant
variables remain fixed.
Current baseline
Experiment ID: ZOYA_PHASE0_BASELINE_00001… See the full description on the dataset page: https://huggingface.co/datasets/tigerking009/zoya-image-1-experiments.abmelt-experiments-exp_20260220_130124indeterminacy-experimentspraxa-behavioral-experiments
Behavioral KV Experiments
Multi-domain protocol and evaluation results for behavioral policy alignment on
COMPASS
scenarios.
Each HF config is one domain bundle: domains/<industry>/<scenario>/ with
frozen split IDs (no query text) plus optional results/<model_key>/<condition>/
run artifacts.
Associated paper
Rosado, E. J. (2026). RFDT: Representation-first decision training for behavioral policy adaptation [Preprint]. URL pending.
Acknowledgments and… See the full description on the dataset page: https://huggingface.co/datasets/EdyVision/praxa-behavioral-experiments.abmelt-experiments-exp_20260218_171120abmelt-experiments-exp_20260220_125440fm-model-experiments-data
FM Model Experiments — Synthetic VLM Training Data (KO + EN)
Annotation data produced while building a native-resolution Korean+English VLM
(GLM-4.6V vision tower transplanted onto a frozen GLM-5.2 743B MoE decoder).
Code + technical report: https://github.com/genonai/fm-model-experiments
This repo contains ANNOTATIONS ONLY (.jsonl). No images are redistributed.
Every row references an image by a relative path (data/...); obtain the images
from the original sources listed below… See the full description on the dataset page: https://huggingface.co/datasets/mncai/fm-model-experiments-data.abmelt-experiments-exp_20260219_182250grok-1-ternary-quant-experiments
Grok-1 SAAQ quantization / route-preservation experiments
Dataset author: Raul Montoya Cardenas (rmems)
SAAQ stands for Spiking Adaptive Activity Quantization, a term coined by
the dataset author.
Attribution: Grok Build: Grok 4.5 (high) packaged the original 2026-08-10
dataset. Codex: GPT-5.6-Sol (OpenAI) implemented, executed, validated, and published the
canonical issue #85 v4 evidence added on 2026-08-24.
Personal research measuring route preservation when packing open… See the full description on the dataset page: https://huggingface.co/datasets/rmems/grok-1-ternary-quant-experiments.abmelt-experiments-exp_20260219_182408abmelt-experiments-exp_20260216_181056mmmlu-bias-experiments
MMMLU Bias Experiments Dataset
Dataset Description
This dataset contains 12 carefully designed experiments to measure language bias and position bias in Large Language Models (LLMs) using multilingual pairwise judgments.
Key Features
12 Experiments: 8 original + 4 position-swapped experiments
11,478 samples per experiment (137,736 total test cases)
Deterministic wrong answers: Uses fixed rule wrong_index = (correct_index + 1) % 4
Perfect correspondence: Wrong… See the full description on the dataset page: https://huggingface.co/datasets/willchow66/mmmlu-bias-experiments.abmelt-experiments-exp_20260215_104855abmelt-experiments-exp_20260216_181641abmelt-experiments-exp_20260215_070527abmelt-experiments-exp_20260215_062019abmelt-experiments-exp_20260215_051916abmelt-experiments-exp_20260215_052622prepared_context_4_experiments_old_train_bert-base-uncasedexperimentsprepared_data_file_zero_shot_prompting_100Q_4_experiments.jsoniatp-experiments
