datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Anime-Background-Finetuning-V1.1
Anime-Background-Finetuning (10143 manually curated by hand images from danbooru and reddit collections)
The dataset contain roughly 2k of anime Screencap data and 8k of scrapped danbooru illustration data.
This is the proccessed version of the dataset meant to be used for my personal finetuning practice project, please visit my RicemanT/Background-Finetuning repo for the raw unprocessed data that you can process yourself.
The dataset have two minor type of processing being done… See the full description on the dataset page: https://huggingface.co/datasets/puruchinera/Anime-Background-Finetuning-V1.1.Anime-Background-Finetuning-V1.1
Anime-Background-Finetuning (10143 manually curated by hand images from danbooru and reddit collections)
The dataset contain roughly 2k of anime Screencap data and 8k of scrapped danbooru illustration data.
This is the proccessed version of the dataset meant to be used for my personal finetuning practice project, please visit my RicemanT/Background-Finetuning repo for the raw unprocessed data that you can process yourself.
The dataset have two minor type of processing being done… See the full description on the dataset page: https://huggingface.co/datasets/RicemanT/Anime-Background-Finetuning-V1.1.Anime-Background-Finetuning-Unprocessed
Anime-Background-Dataset (10143 manually curated by hand images from danbooru and reddit collections)
The dataset contain roughly 2k of Screencap data and 8k of scrapped danbooru illustration data.
It is all raw unprocessed data, the illust folder contain scrapped danbooru tags sidecar .txt on most of the images, while the screencap have non. The processed data is being worked on a seperate repo (Anime-Background-Finetuning)
chorus-backgroundsCAIMAN-ASR-BackgroundNoise
Dataset Card for Myrtle/CAIMAN-ASR-BackgroundNoise
This dataset provides background noise audio, suitable for noise augmentation
while training Myrtle.ai's CAIMAN-ASR models.
Dataset Details
Dataset Description
Curated by: Myrtle.ai
License: Myrtle.ai's modifications to the source data are licensed under
the CC BY 4.0 license.
Some of the original data is under the CC BY 3.0 license; the rest is in the public domain.
Please see the Source Data section… See the full description on the dataset page: https://huggingface.co/datasets/Myrtle/CAIMAN-ASR-BackgroundNoise.Anime-Background-Finetuning-V1.1
Anime-Background-Finetuning (10143 manually curated by hand images from danbooru and reddit collections)
The dataset contain roughly 2k of anime Screencap data and 8k of scrapped danbooru illustration data.
This is the proccessed version of the dataset meant to be used for my personal finetuning practice project, please visit my RicemanT/Background-Finetuning repo for the raw unprocessed data that you can process yourself.
The dataset have two minor type of processing being done… See the full description on the dataset page: https://huggingface.co/datasets/HappyHenAi/Anime-Background-Finetuning-V1.1.openve3m_background_changemusicai-background-music-audio-llm-benchmark
Does Background Music Matter to Speech in Pre-trained Language Models
The completed September 2026 study covers 8 model families, 55 instrumental recordings, and 10 evaluation settings. It studies how adding background music to the same spoken question changes model responses.
Latest release and artifact guide
Technical report PDF
Complete LaTeX project
LaTeX GitHub repository
Matrices, figures, and supporting data
Regenerated speech and mixtures: 550 archives / 250,800… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/musicai-background-music-audio-llm-benchmark.gpn-star-p-uniform-v1-background
marin-dna/gpn-star-p-uniform-v1-background
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the background region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-background.phylop-uniform-v1-background
marin-dna/phylop-uniform-v1-background
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the background region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses the pipeline's pinned phyloP conservation… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-background.vertebrate-v1-background
marin-dna/vertebrate-v1-background
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This
draft covers the background region cohort with all species
scope and preserves source FASTA/2bit letter case.
Anchor eligibility uses the pipeline's pinned phyloP conservation filter.
Sequence case is independent of that filter: lowercase bases preserve source
repeat masking, uppercase bases preserve source… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-background.heao-1-a2-spectrum-backgrounds
HEAO-1 A2 Spectrum Backgrounds
The HEAO-1 A2 spec_back directory serves 1,506 .bck and 36 .alt pointed-phase files from the MED, HED-1 and HED-3 detectors.
Use
from datasets import load_dataset
ds = load_dataset("astro-legacy-archive/heao-1-a2-spectrum-backgrounds", "a2_h1l_1058s084_po.bck", split="train")
row = ds[0]
print(row["CHANNEL"], row["COUNTS"])
1 0
The example configuration is a2_h1l_1058s084_po.bck. Configuration names are the source filename stems.… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/heao-1-a2-spectrum-backgrounds.Caltech101_not_background_test
Dataset Card for "Caltech101_not_background_test"
More Information needed
openve3m_background_change_refcommon_voice_22_yue_w_background_captionMerged JackyHoCL/urban-noise-uganda-61k-caption, OpenSound/AudioCaps
TODO: convert to MP3, reduce size
dota-backgroundrealsense-black-green-background-lerobot
record-immitation-blue-arm-realsense-black-green-backgroundMerged
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
Caltech101_not_background_train
Dataset Card for "Caltech101_not_background_train"
More Information needed
dior-backgroundpi05-libero-plus-background-textures-failures
pi0.5 LIBERO-plus Background Textures Failures
Failure summary artifacts for evaluating TensorAuto/tPi0.5-libero on the LIBERO-plus Background Textures perturbation subset through OpenTau.
Evaluation Setup
Benchmark: LIBERO-plus
Suite: libero_10
Perturbation category: Background Textures
Policy: TensorAuto/tPi0.5-libero
Tasks: 289
Episodes per task: 5
Metric episodes: 1445
Seed schedule: episode seeds 1000 to 1004
Episode length: 520 steps
Source machine: wzxuan… See the full description on the dataset page: https://huggingface.co/datasets/d3d3shan/pi05-libero-plus-background-textures-failures.Background_INCONTEXTXiang_Real_Anime_Background_Relight_Qwen_Edit_2511_Tuned
Relight foreground real person by self-tuned qwen edit 2511 lora, maintain consistency with the style and texture of the anime background.
Market1501-Background-Modified
Dataset Card for "Market1501-Background-Modified"
Dataset Summary
The Market1501-Background-Modified dataset is a variation of the original Market1501 dataset. It focuses on reducing the influence of background information by replacing the backgrounds in the images with solid colors, noise patterns, or other simplified alternatives. This dataset is designed for person re-identification (ReID) tasks, ensuring models learn person-specific features while ignoring background… See the full description on the dataset page: https://huggingface.co/datasets/ideepankarsharma2003/Market1501-Background-Modified.rollout_eval_frazier_55episodes_background_29092026_20261001_150500This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/seriintan/rollout_eval_frazier_55episodes_background_29092026_20261001_150500.coco-backgroundeval_svla_15b_black_backgroundThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 1,
"total_frames": 856,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/shuohsuan/eval_svla_15b_black_background.person-background-dataset
Person-Background Dataset
Diverse people placed in various scene backgrounds. Generated with FLUX.1-dev (people) and FLUX.1-Kontext (background editing).
How It Is Collected
The collect.py script:
Stage 1 – People: Generates 10 diverse people with FLUX.1-dev (different ethnicities, genders, ages).
Stage 2 – Backgrounds: Uses FLUX.1-Kontext img2img to edit each person into scene categories (forest, beach, office, etc.). Preserves person identity while changing only the… See the full description on the dataset page: https://huggingface.co/datasets/nirmalendu01/person-background-dataset.background_400_0818This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"shape": [
7
],
"names": [
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"joint_6.pos"… See the full description on the dataset page: https://huggingface.co/datasets/romalab-cbf/background_400_0818.Yi_Chen_Dancing_White_Background_FramePack_First_Last_Frame_Video_Captioned
Videos Are drived from 以尘动画
Mainly about Genshin-Impact Star-rail and so on.
Thank you very much 🙂
Yi_Chen_Dancing_Animation_Videos_White_Background_Splited_Captioned
