datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rl-game-traces-rise-of-the-tomb-raider
古墓丽影:崛起
This public dataset repository contains gameplay trace data uploaded from F:\古墓丽影:崛起.
Contents
Files: 647
Total local size: 496.11 GB
Generated: 2026-06-06T01:03:39+00:00
File Types
.jsonl: 196
.json: 149
.png: 129
.parquet: 49
.mkv: 49
.txt: 49
.jpg: 15
.exe: 11
Notes
This repository may contain gameplay video, Parquet files, JSON/JSONL metadata, and input event logs.
The license is marked as other; review game footage… See the full description on the dataset page: https://huggingface.co/datasets/yinhuankuang/rl-game-traces-rise-of-the-tomb-raider.raid
🚨 RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors 🚨
🌐 Website, 🖥️ Github, 📝 Paper
RAID is the largest & most comprehensive dataset for evaluating AI-generated text detectors.
It contains over 10 million documents spanning 11 LLMs, 11 genres, 4 decoding strategies, and 12 adversarial attacks.
It is designed to be the go-to location for trustworthy third-party evaluation of both open-source and closed-source generated text detectors.
Load… See the full description on the dataset page: https://huggingface.co/datasets/liamdugan/raid.BraTS-2024-Complete
BraTS 2024 Complete Prepared Dataset
Brain Tumor Segmentation (Leave a like 💖 if this helped you)
Dataset Description
This is an organized and verified version of the BraTS 2024 challenge datasets, including three tumor types.
Included Datasets
Dataset
Type
Cases
Source
BraTS-GLI
Glioma
1,809
Synapse (Dec 2024)
BraTS-MEN-RT
Meningioma + RT
571
Synapse (Feb 2025)
BraTS-PED
Pediatric
348
Cancer Imaging Archive… See the full description on the dataset page: https://huggingface.co/datasets/RaidenShogunUltimate/BraTS-2024-Complete.RadImageNet-VQA
RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering
We introduce RadImageNet-VQA, a large-scale dataset designed for training and benchmarking radiologic VQA on CT and MRI exams. Built from the CT/MRI subset of RadImageNet and its expert-curated anatomical and pathological annotations, RadImageNet-VQA provides 750K images with 7.5M generated samples, including 750K medical captions for visual-text alignment and 6.75M… See the full description on the dataset page: https://huggingface.co/datasets/raidium/RadImageNet-VQA.RAIDThis dataset is for testing the adversarial robustness of AI-Generated Image Detectors, as described in the paper RAID: A Dataset for Testing the Adversarial Robustness of AI-Generated Image Detectors.
rise-of-the-tomb-raider-gameplay-data
古墓丽影:崛起
This public dataset repository contains local gameplay data uploaded from F:\古墓丽影:崛起.
Contents
Files: 433
Total local size: 254.19 GB
Generated: 2026-06-08 19:29:07 UTC
File Types
.jsonl: 136
.json: 105
.png: 102
.mkv: 35
.txt: 33
.parquet: 22
Notes
This repository may contain gameplay video, images, Parquet files, JSON/JSONL metadata, and keyboard/mouse event logs.
The license is marked as other; review game footage… See the full description on the dataset page: https://huggingface.co/datasets/xiaoluo11/rise-of-the-tomb-raider-gameplay-data.Raiden-DeepSeek-R1Click here to support our open-source dataset and model releases!
Raiden-DeepSeek-R1 is a dataset containing creative-reasoning and analytic-reasoning responses, testing the limits of DeepSeek R1's reasoning skills!
This dataset contains:
63k 'creative_content' and 'analytical_reasoning' prompts from microsoft/orca-agentinstruct-1M-v1, with all responses generated by deepseek-ai/DeepSeek-R1.
Responses demonstrate the reasoning capabilities of DeepSeek's 685b parameter R1 reasoning model.… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Raiden-DeepSeek-R1.1920-raider-waite-tarot-public-domainRaid_split1920-raider-waite-tarot-public-domainflaird-raid-pan26SID_Set
Dataset Card for SID_Set
Dataset Summary
We provide Social media Image Detection dataSet (SID-Set), which offers three key advantages:
Extensive volume: Featuring 300K AI-generated/tampered and authentic images with comprehensive annotations.
Broad diversity: Encompassing fully synthetic and tampered images across various classes.
Elevated realism: Including images that are predominantly indistinguishable from genuine ones through mere visual inspection.
Please… See the full description on the dataset page: https://huggingface.co/datasets/RAID-techjam/SID_Set.TacUMI
TacUMI Dataset
Hugging Face: RAIDS666/TacUMI
TacUMI 遥操作采集数据的发布仓库。顶层目录按任务类型划分,每个任务包含从原始视频到训练数据的完整流水线产物。
目录结构
TacUMI/
├── README.md
├── plugin_socket/ # 插插座任务
│ ├── raw_videos/ # 原始 GoPro 视频(约 280 条)及标定素材
│ ├── demos/ # SLAM 流水线中间结果(282 个 demo 目录)
│ ├── dataset_plan.pkl # 数据集规划(位姿对齐、episode 划分等)
│ ├── dataset.zarr.zip # 训练用 replay buffer(4.2… See the full description on the dataset page: https://huggingface.co/datasets/RAIDS666/TacUMI.Genshin_Impact_RaidenShogun_Voice_koreanRAID_none-encoded-gpt2raid-neologism-table-splits
RAID neologism table splits
Source-disjoint RAID train-derived paired splits for the two-token AI detector experiments.
Source dataset: liamdugan/raid, config raid, split train.
Seed: 20260501.
Base source partition: 10000 train source_ids, 3000 test source_ids, overlap 0.
Row format: one human text and one same-source_id AI text per row.
Protocols:
standard_train, standard_test: model, attack, decoding, repetition penalty, and domain sampled randomly.
model_<model>_train… See the full description on the dataset page: https://huggingface.co/datasets/danielfein/raid-neologism-table-splits.testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 12,
"total_frames": 6835,
"total_tasks": 2,
"total_videos": 24,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:12"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/raidavid/test.raiden_garments_folding_baseline_a2
raiden_garments_folding_baseline_a2
LeRobot v2.1 dataset of bimanual YAM (Raiden) garment folding, converted with per-subtask language labels (cell A2).
129 episodes / 95,762 frames / 30 fps
21 unique task strings
Cameras: observation.images.top, left_wrist, right_wrist (224×224, AV1)
State/action: 14-D joints (left 6+gripper, right 6+gripper)
Each LeRobot episode is one labeled subtask (not the high-level collect prompt). Typical sequence per garment:
single out the… See the full description on the dataset page: https://huggingface.co/datasets/Sshawnin/raiden_garments_folding_baseline_a2.raiden_shogun_genshin
Dataset of raiden_shogun/雷電将軍/雷电将军 (Genshin Impact)
This is the dataset of raiden_shogun/雷電将軍/雷电将军 (Genshin Impact), containing 500 images and their tags.
The core tags of this character are long_hair, purple_hair, purple_eyes, breasts, mole, mole_under_eye, large_breasts, hair_ornament, braid, very_long_hair, braided_ponytail, hair_flower, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/raiden_shogun_genshin.Raiden-Mini-DeepSeek-V3.2-SpecialeClick here to support our open-source dataset and model releases!
Raiden-Mini-DeepSeek-V3.2.Speciale is a dataset containing creative-reasoning and analytic-reasoning responses, testing the limits of DeepSeek-V3.2.Speciale's reasoning skills!
This dataset contains:
a default subset of ~8k 'creative_content' and 'analytical_reasoning' prompts from sequelbox/Raiden-DeepSeek-R1, with all responses generated by DeepSeek V3.2 Speciale.
provides an unfiltered look into the reasoning skills of… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Raiden-Mini-DeepSeek-V3.2-Speciale.MetricEval-BodyCT
MetricEval-BodyCT
This repository is a body CT benchmark for evaluating radiology report-generation metrics against
radiologists' judgment.
It covers 100 CT studies (50 chest and 50 abdomen/pelvis), with three candidate reports each. Every
candidate was independently annotated by multiple board-certified radiologists. The reference reports
are de-identified radiology reports from multiple US centers, provided by Segmed and redistributed
under the Data Use Agreement in LICENSE.… See the full description on the dataset page: https://huggingface.co/datasets/raidium/MetricEval-BodyCT.raiden_five_garments_adverserial_rounds1to10_filtered_lay_flatraidex-resultsRaiden-DeepSeek-R1-PREVIEWThis is a preview of the full Raiden-Deepseek-R1 creative and analytical reasoning dataset, containing the first ~6k rows. Get the full dataset here!
This dataset uses synthetic data generated by deepseek-ai/DeepSeek-R1.
The initial release of Raiden uses 'creative_content' and 'analytical_reasoning' prompts from microsoft/orca-agentinstruct-1M-v1.
Dataset has not been reviewed for format or accuracy. All responses are synthetic and provided without editing.
Use as you will.
executable-counterfactuals
Introduction
This repo contains all training and evaluation datasets used in "Executable Counterfactuals: Improving LLMs' Causal Reasoning Through Code". This work has been published in ICLR 2026.
Arxiv Paper
Github Repo (Work in Progress)
Counterfactual reasoning, a hallmark of intelligence, consists of three steps: inferring latent variables from observations (abduction), constructing alternative situations (interventions), and predicting the outcomes of the alternatives… See the full description on the dataset page: https://huggingface.co/datasets/Raidriar-Dai/executable-counterfactuals.raiden_five_garments_baseline_rounds1to10_filtered_lay_flatbackuplienchiangprobslistraiden_single_garment_lay_flat_blue_day2
raiden_single_garment_lay_flat_blue_day2
Oct 8, 2026 pants lay-flat on the YAM / Raiden. Raw episodes 0000–0177. 178 episodes, 183193 frames, 30 fps, about 101.8 min.
Task string: lay it flat
High-level prompt: lay the garment flat
Robot: yam
Included takes to review before training: 0001 and 0154 were still pending in the console, and 0165 has a camera/robot stall (coverage 0.97, one kept piece).
raiden_single_garment_lay_flat_blue
raiden_single_garment_lay_flat_blue
Oct 7, 2026 pants lay-flat on the YAM / Raiden. 181 episodes, 182240 frames, 30 fps, about 101.2 min.
Task string: lay it flat
High-level prompt: lay the garment flat
Robot: yam
raiden_five_garments_adverserial_rounds1to10_singleout_layflat
