datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Beta-Pre-Train-Corpus
Reactive AI / Beta Pre-Train Corpus
Pre-training corpus for RxT-Beta models, created from public & open datasets. Includes high-quality english and polish web crawl data, mathematic and scientific subsets,
and code in different programming languages.
2k subsets are filtered for 1024-2048 tokens, except MegaMath Web Pro and GitHub Code subsets, that were filtered for 512-2048 tokens
Subsets & original datasets
FineWeb-Edu
fineweb-edu-s100 (51.3M examples) - 50% of… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/Beta-Pre-Train-Corpus.Beta-Hybrid-Interaction-SFTReact
React — Multi-Task Tactile-Visual Manipulation
Dense, contact-rich, synchronized multimodal interaction data collected from human hands holding handheld GelSight tactile sensors (no robot arm). Intended for tactile-visual dynamics / world-model learning.
133 min · 240 k frames @ 30 Hz · 3× RGB + 2× GelSight + OptiTrack · 2 tasks
Format — LeRobot-style video release
Each episode ships as 5 MP4 video streams (640×480, H.264) + a per-frame parquet of poses and… See the full description on the dataset page: https://huggingface.co/datasets/yxma/React.ord-data
ord-data
Getting the Data
The datasets live under data/ and are stored with
Git LFS. LFS reads are redirected to the
Hugging Face mirror
via .lfsconfig, so dataset objects are fetched from Hugging
Face's CDN rather than from GitHub's shared (and limited) LFS bandwidth. This is
automatic — you do not need to configure anything.
Option 1: Clone the repository
git clone https://github.com/open-reaction-database/ord-data.git
With Git LFS installed… See the full description on the dataset page: https://huggingface.co/datasets/open-reaction-database/ord-data.ReactiveGWM-Datasets
ReactiveGWM-Datasets: Strategy-Aligned Rollouts for Reactive Game World Models
📚 Datasets-Introduction
ReactiveGWM-Datasets is the strategy-aligned training corpus that powers
ReactiveGWM, a game world
model that decouples player control from NPC autonomy. To learn that
decoupling, the model needs supervision that pairs each gameplay clip with
both a per-frame action stream (what the player did) and a high-level
NPC description (what the NPC tried to do, and under… See the full description on the dataset page: https://huggingface.co/datasets/INV-WZQ/ReactiveGWM-Datasets.webui-react-htmlcssjs-8740
WebUI React + HTML/CSS/JS 8,740
Curated export from ronantakizawa/webui containing every row where framework = react, plus 4,000 additional rows where framework = vanilla.
Screenshots: 8,740
React / vanilla HTML-CSS-JS rows: 4,740 / 4,000
Unique sample IDs: 2,914
Train / validation / test: 7,501 / 456 / 783
Viewports: 2,914 desktop / 2,913 mobile / 2,913 tablet
Images are stored as real image files and verified with Pillow.
viewer.html is a self-contained, paginated local… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/webui-react-htmlcssjs-8740.smol-smoltalk-Interaction-SFT
Dataset Card for ReactiveAI/Smol-Smoltalk Interaction SFT
Derived from HuggingFaceTB/smol-smoltalk. Made for Interaction Supervised Fine-Tuning of Reactive Transformer
Proof-of-Concept models, especially RxT-Beta.
Dataset Details
Dataset Description
Reactive Transformers are processing only the single interactions in real-time and using Short-Term Memory to store information from previous interactions.
Before the model is able to use it's memory, it has to be… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/smol-smoltalk-Interaction-SFT.Octant_CYP_inhibition_reactivity_blog_release
OpenADMET Octant CYP Inhibition & Reactivity
Data release from the OpenADMET consortium, generated by Octant Bio.
This dataset accompanies the blog post Building the OpenADMET Data Engine.
Source code, assay protocols, and raw TSV files are on GitHub.
Overview
Cytochrome P450 (CYP) enzymes drive the oxidative metabolism of most drugs and are a primary cause of drug-drug interactions (DDIs).
Despite their importance, public CYP datasets are sparse, noisy, and… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/Octant_CYP_inhibition_reactivity_blog_release.BioDEX-Reactions
Dataset Card for "BioDEX-Reactions"
More Information needed
hle-react
AggAgent ReAct Rollouts - HLE
Dataset Description
AggAgent is an agentic aggregation framework that scales long-horizon agents at test time by sampling multiple parallel rollouts from a base agent and then aggregating their evidence and solutions. This dataset card releases the ReAct base rollouts that AggAgent consumes, i.e. single-agent trajectories produced before any aggregation step.
Each rollout was generated by running a ReAct-style deep-research… See the full description on the dataset page: https://huggingface.co/datasets/yoonsanglee/hle-react.algebraic-stack-fixedpzhrd-programmable-zeno-holonomic-reaction-darkspace
PZHRD — Programmable Zeno–Holonomic Reaction Darkspace
Tangent-matched recovery, geometric reaction addressing, deferred-commit logical chemistry, and error-corrected matter construction
Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiRelease: v1.0.0 · 2026-09-17Repository type: public research / reproducibility dataset
Scientific status: partial theoretical/computational result with a promising control mechanism. This release does not demonstrate a universal… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/pzhrd-programmable-zeno-holonomic-reaction-darkspace.Beta-Code
Reactive AI / Beta Code
Code-based pre-training corpus for RxT-Beta models, created from public & open datasets. Includes code in different programming languages.
Subsets are divided into short (< ~1024 tokens) and long (> ~1024 tokens) categories.
Original dataset
It's created from codeparrot datasets:
Python subsets from codeparrot/codeparrot-clean
other subsets from codeparrot/github-code-clean
react-code-instructions
React Code Instructions
Popular Queries
Number of instructions by Model
Unnested Messages
Instructions Added Per Day
Dataset of Claude Artifact esque React Apps generated by Llama 3.1 70B, Llama 3.1 405B, and Deepseek Chat V3.
Examples
Virtual Fitness Trainer Website
LinkedIn Clone
iPhone Calculator
Chipotle Waitlist
Apple Store
RxQ-SMATfinepdfs-edu-betaTDC_skin_reaction
TDC Skin Reaction
Skin Reaction dataset dataset [1], part of TDC [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict whether the drug can cause immune reaction that leads to skin sensitization.
Characteristic
Description
Tasks
1
Task type
classification
Total samples
404
Recommended split
scaffold
Recommended metricAUROC
References
[1]
Alves, Vinicius M., et al.
"Predicting chemically-induced skin… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/TDC_skin_reaction.ReactionSmiles
Reaction SMILES Dataset
A collated dataset of 3.4M unique chemical reaction SMILES strings compiled from multiple public sources for use in pre-training and fine-tuning chemical language models.
Reaction SMILES (Simplified Molecular Input Line Entry System) extend the standard SMILES notation to represent complete chemical reactions. They encode reactants, reagents/catalysts, and products in a single text string using the > delimiter:
reactants>reagents>products
For example:… See the full description on the dataset page: https://huggingface.co/datasets/Derify/ReactionSmiles.RxQ-iSFTReactiveGWM-v2-Datasets
ReactiveGWM v2 Datasets
This repository contains the selected HNM and Street Fighter III: New
Generation (SF3) datasets. The original FFV1/MKV videos, first frames, masks,
native annotations, action records, indices, and prepared VAE/T5 caches are
stored without compression in tar volumes of approximately 2 GiB. Files
retain their original formats and paths inside the tar archives. Game-generation
code, ROMs, BIOS files, and runtime binaries are not included.
Subset
Train… See the full description on the dataset page: https://huggingface.co/datasets/INV-WZQ/ReactiveGWM-v2-Datasets.ConversationalRetrieval-SMAT
ReactiveAI / ConverstationalRetrieval Dataset for Supervised Memory-Aware Training (SMAT)
Description in progress
beta-reasoningsd-reactor-node
ReActor Node 0.1.1b for ComfyUI
The Fast and Simple "roop-like" Face Swap Extension Node for ComfyUI, based on ReActor (ex Roop-GE) SD-WebUI Face Swap Extension
This Node goes without NSFW filter (uncensored, use it on your own responsibility)
Disclaimer | Installation | Usage | Troubleshooting | Updating
Disclaimer
This software is meant to be a productive contribution to the rapidly growing AI-generated media industry. It will help artists with tasks… See the full description on the dataset page: https://huggingface.co/datasets/crystantine/sd-reactor-node.uniprot_reactions
Dataset Details
Dataset Description
Protein sequences and the reactions these can catalyze.
Curated by:
License: MIT
Dataset Sources
data source
Citation
BibTeX:
@article{10.1093/nar/gkac1052,
author = {The UniProt Consortium},
title = {UniProt - the Universal Protein Knowledgebase in 2023},
journal = {Nucleic Acids Research},
volume = {51},
number = {D1},
pages = {D523-D531},
year = {2022},
month = {11},
issn = {0305-1048},
doi =… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/uniprot_reactions.btc-news-forward-price-reactions
BTC News and Price Reactions
120,981 multilingual news events joined to weak semantic
annotations and BTC returns before and after each observation, from 1 minute
to 24 hours. The release is designed for event studies, classifier
bootstrapping, temporal evaluation, and research on how observed news aligns
with market movement.
Start here: licensed_news_price_reaction is the analysis-ready view with
attributed reusable titles/RSS summaries, weak labels, entry prices, and all
six… See the full description on the dataset page: https://huggingface.co/datasets/dymyt-ry/btc-news-forward-price-reactions.ICU-REACT
ICU-REACT
ICU-REACT is a clinician-supervised dataset for clinical reasoning and information retrieval in the intensive care unit (ICU), developed for fine-tuning and benchmarking large language models (LLMs).
ICU-REACT was constructed using a clinician-in-the-loop annotation framework designed to capture how clinicians identify relevant patient information and integrate it into diagnostic and treatment decisions. The dataset includes a clinician-refined seed training set, a… See the full description on the dataset page: https://huggingface.co/datasets/iheallab/ICU-REACT.react-llama
The ReAct Llama Dataset
Dataset Summary
This dataset contains 3,538 correct ReAct trajectories generated using llama2-70b (Q5_K_M quant).
It follows the format used in the ReAct paper.ReAct trajectories were generated using a modified version of the hotpotqa.ipynb file from the ReAct repo.
The model was prompted in the following format (5-shot) to generate these traces:
Solve a question answering task with interleaving Thought, Action, Observation steps. Thought can… See the full description on the dataset page: https://huggingface.co/datasets/xz56/react-llama.Beta-Hybrid-SMAT
Reactive AI / Beta Hybrid SMAT
Multi-turn conversational dataset with hybrid reasoning for Supervised Memory Aware Training (SMAT) of Reactive Transformer MVP Beta models
Flame-Waterfall-React
Flame-Waterfall-React: A Structured Data Synthesis Dataset for Multimodal React Code Generation
Flame-Waterfall-React is a dataset synthesized using the Waterfall-Model-Based Synthesis method, Advancing Vision-Language Models in Front-End Development via Data Synthesis. This dataset is designed to train vision-language models (VLMs) for React code generation from UI design mockups and specifications.
The Waterfall synthesis approach mimics real-world software development by… See the full description on the dataset page: https://huggingface.co/datasets/Flame-Code-VLM/Flame-Waterfall-React.NVIDIA-Nemotron-IF-Chat-v2-rx
