datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bootstrap-latent-thought-dataThis dataset is associated with the paper Reasoning to Learn from Latent Thoughts. It contains data used for pretraining language models with a focus on improving data efficiency by modeling and inferring latent thoughts underlying the text generation process, such as on reasoning-intensive math corpus. An expectation-maximization algorithm is developed for models to self-improve their self-generated thoughts and data efficiency.
synth-bootstrap-trialbootstrapvue-standardSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: BootstrapVue 2.23
Documentation Data Source Link: https://bootstrap-vue.org/docs/
Data Source License: https://github.com/bootstrap-vue/bootstrap-vue/blob/dev/LICENSE
Data Source Authors: BootstrapVue Team
AI Benchmarks by Data Agents © 2025 RELAI.AI · Licensed under CC BY 4.0. Source: https://relai.ai
bootstrap_sms_v2_repeat_1hgt-bootstrap-v1-synthetic
HGT Bootstrap V1 Synthetic Pairs
8,112 synthetic mutation pairs generated via ICI-DC (Interleaved Codon Insertion — Double Consensus) using the SAD coeff1.5 checkpoint as Model A and HyenaDNA-tiny-1k as Model B.
Generation Details
Model A: Nhoodie/omni-dna-sad-mutation (SAD coeff1.5 checkpoint)
Model B: LongSafari/hyenadna-tiny-1k-seqlen-hf (Legacy DC)
Source sequences: 1,014 unique sequences from 8 taxonomic domains
Gap intervals: [3, 4, 5, 6] (codon distance between… See the full description on the dataset page: https://huggingface.co/datasets/Nhoodie/hgt-bootstrap-v1-synthetic.ppl-synthesis-sft-bootstrap
SynthStats PPL Synthesis SFT Bootstrap
This dataset contains natural-language modelling prompts paired with
probabilistic programs, written in the probabilistic programming
languages PyMC (Python) and LazyPPL (Haskell), for supervised
fine-tuning (SFT).
Each row has these fields:
prompt: natural-language modelling task.
reasoning_trace: modelling rationale for the program.
completion: one fenced program block.
complexity: coarse task complexity label.
metadata: runtime… See the full description on the dataset page: https://huggingface.co/datasets/SynthStats/ppl-synthesis-sft-bootstrap.AcitvityNet-Captions-bootstrapped-5Khgt-bootstrap-v2-synthetic
HGT Bootstrap V2 Synthetic Pairs
259,896 synthetic mutation pairs (129,948 train + 129,948 eval) generated via ICI-DC using the bootstrap S2 checkpoint.
Generation Details
Model A: Nhoodie/omni-dna-sad-mutation-bootstrap (S2 bootstrap checkpoint)
Model B: LongSafari/hyenadna-tiny-1k-seqlen-hf
Source sequences: 4,998 unique sequences from 8 taxonomic domains
Gap intervals: [3, 4, 5, 6] (codon distance between gaps)
Seeds: Train=[42, 137, 7, 23, 31, 89, 53], Eval=[101, 157… See the full description on the dataset page: https://huggingface.co/datasets/Nhoodie/hgt-bootstrap-v2-synthetic.Bootstrap5bootstrap_oai_ptmodel-cards-ml-metadata-bootstrap
davanstrien/model-cards-ml-metadata-bootstrap
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over librarian-bots/model_cards_with_metadata.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
librarian-bots/model_cards_with_metadata (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
base model name, context length, training method, training dataset name, benchmark name… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/model-cards-ml-metadata-bootstrap.training-methods-bootstrap
davanstrien/training-methods-bootstrap
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over /input/cleaned-cards.parquet.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
/input/cleaned-cards.parquet (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
training method
Confidence threshold
0.7
Samples processed
10000
Total entities extracted
4278
Inference device
cuda… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/training-methods-bootstrap.bootstrap_sms
Dataset Card for "bootstrap_sms"
More Information needed
curated-v2-combined-dpo-bootstrapeval-mentions-bootstrap-v2
davanstrien/eval-mentions-bootstrap-v2
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over /input/cleaned-cards-quality.parquet.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
/input/cleaned-cards-quality.parquet (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
benchmark name, evaluation metric
Confidence threshold
0.6
Samples processed
5000
Total entities… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/eval-mentions-bootstrap-v2.grpo-5-sft-bootstrapeval-mentions-bootstrap
davanstrien/eval-mentions-bootstrap
Bootstrap NER dataset produced by urchade/gliner_multi-v2.1 over /input/cleaned-cards.parquet.
Generated using uv-scripts/gliner/extract-entities.py.
Provenance
Source dataset
/input/cleaned-cards.parquet (split train)
Text column
card
Bootstrap model
urchade/gliner_multi-v2.1
Entity types
benchmark name, evaluation dataset, evaluation metric
Confidence threshold
0.6
Samples processed
10000
Total entities… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/eval-mentions-bootstrap.tarsier-bootstrapperbootstrap_oai_pt_thinkbootstrap_agreement_long_17grpo-5-sft-bootstrap-qwen3-4b-thinking-2507
BRlkl/grpo-5-sft-bootstrap-thinking
Derived from BRlkl/grpo-5-sft-bootstrap-qwen3-4b-thinking-2507.
This version repairs the plan column only for rows where blacklisted = true.
Transformation:
Parse the plan JSON.
Read walk[0].
If walk[0] is the wrapped prompt form:
CONVERSATION_HISTORY: [Empty] ... Generate {"walk":[...]} for NEW_USER_MESSAGE.
then replace it with just the embedded NEW_USER_MESSAGE text.
Leave all non-blacklisted rows unchanged.
Audit summary:
Total rows:… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/grpo-5-sft-bootstrap-qwen3-4b-thinking-2507.sam3-ls-bootstrap-demo
davanstrien/sam3-ls-bootstrap-demo
Bootstrap dataset produced by running facebook/sam3 over a small set of test images and storing the predictions in a Label Studio project for review.
This is a proof-of-concept artifact demonstrating an end-to-end "unlabeled images → bootstrapped dataset" workflow on Hugging Face infrastructure. The predictions in this dataset are SAM3 outputs — not human-reviewed.
Workflow
Images imported into Label Studio project 20 on… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/sam3-ls-bootstrap-demo.bootstrap_leslie_long_93bootstrap_leslie_long_80bootstrap_leslie_long_94bootstrap_prevalence_long_10bootstrap_prevalence_long_18bootstrap_agreement_long_22bootstrap_agreement_long_43bootstrap_sms_v2
