datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
graphical-bootstrap-correlator-dataset
Graphical Bootstrap Correlator Dataset
This dataset contains large-scale graph-structured data arising from high-order perturbative computations of four-point correlators in planar $\mathcal{N}=4$ super Yang--Mills theory.
The data consists of denominator graphs (d-graphs) appearing in the graphical bootstrap formulation of correlators. Each graph is associated with a binary label indicating whether it contributes to the correlator at a given perturbative order.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Gabriele-dian/graphical-bootstrap-correlator-dataset.bootstrap-latent-thought-dataThis dataset is associated with the paper Reasoning to Learn from Latent Thoughts. It contains data used for pretraining language models with a focus on improving data efficiency by modeling and inferring latent thoughts underlying the text generation process, such as on reasoning-intensive math corpus. An expectation-maximization algorithm is developed for models to self-improve their self-generated thoughts and data efficiency.
synth-bootstrap-trialgraphical-bootstrap-correlator-dataset
Graphical Bootstrap Correlator Dataset
This dataset contains large-scale graph-structured data arising from high-order perturbative computations of four-point correlators in planar $\mathcal{N}=4$ super Yang--Mills theory.
The data consists of denominator graphs (d-graphs) appearing in the graphical bootstrap formulation of correlators. Each graph is associated with a binary label indicating whether it contributes to the correlator at a given perturbative order.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/anonymous314/graphical-bootstrap-correlator-dataset.maskscore-rung-1-bootstrap
MaskScore Rung 1 — Bootstrap (5 of 8 stubs)
Walking-skeleton implementation of MaskScore Rung 1. Five of the eight MASKSCORE.md
stubs are filled with real content from a synthetic ANNY bootstrap (rest pose + rank1
identity + rank5 perturbation). Text, Speech, and Video stubs are deferred to Rung 2 —
the bootstrap has no transcript, no audio, and no video, and CLAUDE.md's ETNF rule
forbids putting a null in for the missing input.
Each stub ships as three ZSTD-compressed parquets:… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/maskscore-rung-1-bootstrap.guile-bootstrapbootstrapvue-standardSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: BootstrapVue 2.23
Documentation Data Source Link: https://bootstrap-vue.org/docs/
Data Source License: https://github.com/bootstrap-vue/bootstrap-vue/blob/dev/LICENSE
Data Source Authors: BootstrapVue Team
AI Benchmarks by Data Agents © 2025 RELAI.AI · Licensed under CC BY 4.0. Source: https://relai.ai
bootstrap_sms_v2_repeat_1ppl-synthesis-sft-bootstrap
SynthStats PPL Synthesis SFT Bootstrap
This dataset contains natural-language modelling prompts paired with
probabilistic programs, written in the probabilistic programming
languages PyMC (Python) and LazyPPL (Haskell), for supervised
fine-tuning (SFT).
Each row has these fields:
prompt: natural-language modelling task.
reasoning_trace: modelling rationale for the program.
completion: one fenced program block.
complexity: coarse task complexity label.
metadata: runtime… See the full description on the dataset page: https://huggingface.co/datasets/SynthStats/ppl-synthesis-sft-bootstrap.HTML-CSS-BOOTSTRAP-JAVASCRIPThexo-bootstrap-corpus
Hexo Human Corpus
Encoding-free corpus of decisive human Hex Tac Toe games — hexagonal grid,
six-in-a-row to win (player 1 opens with 1 move, then both players play 2 moves
per turn; the board is theoretically infinite).
Each line is one game as a raw axial move list + outcome. Nothing about any
neural-network encoding is baked in — no planes, no fixed board size, no action
space. Read it with the stdlib json module and build whatever representation
you want.
Files… See the full description on the dataset page: https://huggingface.co/datasets/timmyburn/hexo-bootstrap-corpus.hexo-bootstrap-corpus
Hexo Human Corpus
Encoding-free corpus of 6,902 decisive human Hex Tac Toe games — hexagonal
grid, six-in-a-row to win (player 1 opens with 1 move, then both players play 2
moves per turn; the board is theoretically infinite).
Each line is one game as a raw axial move list + outcome. Nothing about any
neural-network encoding is baked in — no planes, no fixed board size, no action
space. Read it with the stdlib json module and build whatever representation
you want.… See the full description on the dataset page: https://huggingface.co/datasets/Yulolam/hexo-bootstrap-corpus.hgt-bootstrap-v1-synthetic
HGT Bootstrap V1 Synthetic Pairs
8,112 synthetic mutation pairs generated via ICI-DC (Interleaved Codon Insertion — Double Consensus) using the SAD coeff1.5 checkpoint as Model A and HyenaDNA-tiny-1k as Model B.
Generation Details
Model A: Nhoodie/omni-dna-sad-mutation (SAD coeff1.5 checkpoint)
Model B: LongSafari/hyenadna-tiny-1k-seqlen-hf (Legacy DC)
Source sequences: 1,014 unique sequences from 8 taxonomic domains
Gap intervals: [3, 4, 5, 6] (codon distance between… See the full description on the dataset page: https://huggingface.co/datasets/Nhoodie/hgt-bootstrap-v1-synthetic.drivevla-w0-ar-bootstrap
DriveVLA-W0 NAVSIM AR Action Expert bootstrap
This directory is a server-side bootstrap bundle, not a model or dataset
mirror. It contains reproducible instructions, pinned-source metadata, server
download/verification code, training profiles, and local static-check reports.
Local-stage boundary
The local stage was run through WSL2 Ubuntu-22.04. The Linux source checkout,
revision lock, source archive, source audit, and CPU-only static smoke test have
completed.… See the full description on the dataset page: https://huggingface.co/datasets/wannac1/drivevla-w0-ar-bootstrap.hgt-bootstrap-v2-synthetic
HGT Bootstrap V2 Synthetic Pairs
259,896 synthetic mutation pairs (129,948 train + 129,948 eval) generated via ICI-DC using the bootstrap S2 checkpoint.
Generation Details
Model A: Nhoodie/omni-dna-sad-mutation-bootstrap (S2 bootstrap checkpoint)
Model B: LongSafari/hyenadna-tiny-1k-seqlen-hf
Source sequences: 4,998 unique sequences from 8 taxonomic domains
Gap intervals: [3, 4, 5, 6] (codon distance between gaps)
Seeds: Train=[42, 137, 7, 23, 31, 89, 53], Eval=[101, 157… See the full description on the dataset page: https://huggingface.co/datasets/Nhoodie/hgt-bootstrap-v2-synthetic.model-cards-ml-metadata-bootstrap
davanstrien/model-cards-ml-metadata-bootstrap
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over librarian-bots/model_cards_with_metadata.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
librarian-bots/model_cards_with_metadata (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
base model name, context length, training method, training dataset name, benchmark name… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/model-cards-ml-metadata-bootstrap.bootstrap_sms
Dataset Card for "bootstrap_sms"
More Information needed
Bootstrap5bootstrap_oai_ptLlama2-7B-generic-predictions-starwars-bootstrap-coef-10-oncecurated-v2-combined-dpo-bootstrapjade-propainter-bootstraptraining-methods-bootstrap
davanstrien/training-methods-bootstrap
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over /input/cleaned-cards.parquet.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
/input/cleaned-cards.parquet (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
training method
Confidence threshold
0.7
Samples processed
10000
Total entities extracted
4278
Inference device
cuda… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/training-methods-bootstrap.eval-mentions-bootstrap-v2
davanstrien/eval-mentions-bootstrap-v2
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over /input/cleaned-cards-quality.parquet.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
/input/cleaned-cards-quality.parquet (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
benchmark name, evaluation metric
Confidence threshold
0.6
Samples processed
5000
Total entities… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/eval-mentions-bootstrap-v2.AcitvityNet-Captions-bootstrapped-5Kgrpo-5-sft-bootstraprunpod-minimax-h3-bootstraptarsier-bootstrapperLlama2-7B-generic-predictions-starwars-bootstrap-coef-10-once-augmentedbootstrap_agreement_long_17
