datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
parity-experiments
Upload your Adapter Oracle and Parity results
This dataset saves the oracle and parity experiment logs for adapters. Please upload them according to the following format and draft a PR.
adapters/
└── {adapter_name}/
├── README.md # Results overview, directory structure, trajectory interpretation, notes, etc. This should be DIFFERENT than the adapter REAMDE.
├── config.yaml # The yaml file that can be directly used to run parity experiments in Harbor.
├──… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/parity-experiments.RRC_Experimentshqnet-experimentsjax-fli-experiments
jax-fli experiments
Data, samples, and reference catalogs for the jax-fli
forward-modelling experiments. Each experiment is exposed as one or more
HuggingFace dataset configs; load a config with datasets.load_dataset.
This dataset holds the accuracy experiments and feeds the Results Explorer. The scaling benchmarks are in ASKabalan/jax-fli-scaling, and the MAP and chain outputs in ASKabalan/jax-fli-sampling.
Experiment 00 — CosmoGrid reference
A single CosmoGrid… See the full description on the dataset page: https://huggingface.co/datasets/ASKabalan/jax-fli-experiments.experimentsdl2l-experiments
DL2L Experiments Dataset
Simulation trajectory data from the DL2L
distributed artificial life simulator, used to train JEPA world models.
See felipedreis/dl2l-jepa for the trained models.
Dataset structure
Data is organized by experiment prefix. Each prefix contains parquet files for
model training and a stats.json with dataset metadata.
p9/
train.parquet # single-encoder training set (trials 1–8)
val.parquet # single-encoder validation set… See the full description on the dataset page: https://huggingface.co/datasets/felipedreis/dl2l-experiments.experimentsatomic-metrics-experiments
Atomic Metrics Experiment Artifacts
Run directories from
Atomic Metrics, including
extracted metric banks, generated domain prompts, batch/refine snapshots, and
BT/LR eval summaries.
Logs are omitted. API keys are not included; scoring used environment
credentials at runtime.
pisa-experiments
Pisa Experiments
This repository contains the PisaBench, training data, model checkpoints, introduced in PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop.
PisaBench
Real World Videos
We curate a dataset comprising 361 videos demonstrating the dropping task.Each video begins with an object suspended by an invisible wire in the first frame. We cut the video clips to begin as soon as the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/pisa-experiments.gemma-crafter-five-experiments-20260915bert-mlm-experiments-en
Unified English MLM Pre-training Corpus (80M Rows)
This dataset is a massive, diverse, multi-domain English text corpus explicitly engineered for pre-training and domain-adaptation of BERT-style models via Masked Language Modeling (MLM). It aggregates over 80 million rows of text, completely stripped of auxiliary metadata, labels, and identifiers to expose purely raw text strings.
Dataset Details
Repository ID: 8Opt/bert-mlm-experiments-en
Total Rows: 80,489,226… See the full description on the dataset page: https://huggingface.co/datasets/LakoreAI/bert-mlm-experiments-en.transformers-merge-experimentskomorebi-painter-experimentsexperiment-speaker-embeddingLeRing_JFM_experiments
Overview
This dataset repository contains training data and experimental recordings to measure the rotations of particles suspended in viscous shear flows.
This dataset is intended to support the development of machine learning research in fluid dynamics, especially in the study of multi-phase flows.
Content
The repository contains training data and extensive experimental measurements of single particle suspended in confined shear flows in the viscous and small-inertial… See the full description on the dataset page: https://huggingface.co/datasets/ddg93/LeRing_JFM_experiments.hle-flowbench-experiments-20260829
HLE FlowBench experiment archive
Private migration snapshot of the local HLE text-only 100-question research
program through 2026-09-01. It preserves the formal and smoke runs, per-question
Codex/Claude/Kimi sessions, workflow attempts and metrics, evaluator state,
scores, monitoring, experiment controllers, reports, analyses, the paused-run
migration package, source Git bundles, and HLE-related host orchestration
sessions.
The current Chinese experiment status, validity… See the full description on the dataset page: https://huggingface.co/datasets/Changyeli03/hle-flowbench-experiments-20260829.fine-tuning-experiments-082023evalexplorer-classify-experiments
EvalExplorer document classifier: experiments
The question
When an evaluation report enters EvalExplorer, the ingestion pipeline sends its first pages to a large LLM
(gpt-oss-120b, with Gemini 2.5 Flash and Qwen 3 235B as fallbacks), which returns five labels: evaluation
approach (mixed methods, experimental, ...), type (impact evaluation, systematic review, ...), timing
(baseline, midterm, endline), themes (global health, governance, ...) and countries (ISO… See the full description on the dataset page: https://huggingface.co/datasets/baobabtech/evalexplorer-classify-experiments.BindCap-F2L-ExperimentsSciGA-for-experiments-hfcibench-experiments
CIBench Experiments
Reproducibility packages for CIBench — the stateless, replayable benchmark engine for the 1M–10M token long-context era.
If a benchmark result cannot be replayed from its manifest alone, it did not happen.
Every sub-directory in this dataset is a self-contained experiment package: per-run manifests, content-addressed canonical JSON, ResultRecord with full scoring + signed provenance, per-item OpenTelemetry gen_ai_* call metrics, retrieved evidence, a… See the full description on the dataset page: https://huggingface.co/datasets/publicus-ai/cibench-experiments.soft-prompt-experiments-archive-20260918
Soft prompt 实验归档
用于查阅和恢复的历史研究记录,涵盖数学与代码任务。共 45 个运行目录,包含教师生成数据、评测输出、原始配置和已有 prompt 检查点。部分目录仅有评测、复核或失败记录,不能将目录数量理解为成功实验数量。
快速查阅
实验总览:模型系列、任务、规模与记录状态。
CSV 索引 / JSON 索引:便于筛选和定位。
archives/:按实验分别压缩的原始文件。
manifests/:各文件 SHA-256 与归档路径。
系列包括 AReaL Boba2、GPT-OSS/Swallow、MiMo、X-Coder/Qwen3、Nemotron、Klear、Mellum2、OLMo3、Polaris 和 Poro2。页面不展开具体方法或实现细节;原始配置仍保留在归档内供恢复。
状态与注意事项
上传完成以 ARCHIVE_COMPLETE.json 为准;文件不存在时表示仍在上传。 每个归档都经过完整下载的 SHA-256 校验。… See the full description on the dataset page: https://huggingface.co/datasets/namezz/soft-prompt-experiments-archive-20260918.synthetic_experimentspeacock-data-public-datasets-experimentsgrpo-gsm8k-experimentsself-organisation-experiments
Self-Organisation Experiment Data
Experiment data for the self-organisation project investigating whether self-replication can emerge spontaneously in continuous neural network weight space.
Result: Negative — continuous weight space appears to lack the computational primitives required for spontaneous self-replication. See the code repository for full analysis.
Structure
autoresearch/ — Automated research runs across 11 experiment configurations
phase1*/ — BFF… See the full description on the dataset page: https://huggingface.co/datasets/LindaP/self-organisation-experiments.jbcs2025_experiments_report
JBCS 2025: Experimental Artefacts for AES in Brazilian Portuguese
This repository contains all experimental artefacts (logs, configurations, predictions, and evaluation results) described in the paper:
Exploring the Usage of LLMs for Automatic Essay Scoring in Brazilian Portuguese EssaysAndré Barbosa, Igor Cataneo Silveira, Denis Deratani MauáTODO
📦 What's in this dataset repo?
This dataset is not a training dataset. Instead, it provides comprehensive logs and… See the full description on the dataset page: https://huggingface.co/datasets/kamel-usp/jbcs2025_experiments_report.thesis-experiments-dataMedStyleAudit-Experiments
MedStyleAudit Experiments
This repository contains experimental artifacts, metadata, intermediate results, and evaluation outputs associated with MedStyleAudit, a framework for detecting acquisition-style shortcuts in medical imaging models.
The repository is intended to support the reproducibility and transparency of the experiments reported in the MedStyleAudit study.
Overview
Medical imaging models may exploit acquisition-related, institution-specific, or… See the full description on the dataset page: https://huggingface.co/datasets/PeiyuanHao/MedStyleAudit-Experiments.language-decoded-experiments
Language Decoded — Experiment Tracking
Central hub for training logs, configurations, evaluation results, and analysis for the Language Decoded project. The project originated as a proposal during Cohere's Tiny Aya Expedition (March 2026 hackathon) and was extended into Phase 3 for the accompanying paper.
Submitted paper title (2026-05-26): Language, Decoded: Exploring the Impact of Fine-Tuning a Multilingual Model on Native-Language Code
⚠️ Phase 3 numbers — read… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-experiments.
