datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Beta-Pre-Train-Corpus
Reactive AI / Beta Pre-Train Corpus
Pre-training corpus for RxT-Beta models, created from public & open datasets. Includes high-quality english and polish web crawl data, mathematic and scientific subsets,
and code in different programming languages.
2k subsets are filtered for 1024-2048 tokens, except MegaMath Web Pro and GitHub Code subsets, that were filtered for 512-2048 tokens
Subsets & original datasets
FineWeb-Edu
fineweb-edu-s100 (51.3M examples) - 50% of… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/Beta-Pre-Train-Corpus.Beta-Hybrid-Interaction-SFTReact
React — Multi-Task Tactile-Visual Manipulation
Dense, contact-rich, synchronized multimodal interaction data collected from human hands holding handheld GelSight tactile sensors (no robot arm). Intended for tactile-visual dynamics / world-model learning.
133 min · 240 k frames @ 30 Hz · 3× RGB + 2× GelSight + OptiTrack · 2 tasks
Format — LeRobot-style video release
Each episode ships as 5 MP4 video streams (640×480, H.264) + a per-frame parquet of poses and… See the full description on the dataset page: https://huggingface.co/datasets/yxma/React.ord-data
ord-data
Getting the Data
The datasets live under data/ and are stored with
Git LFS. LFS reads are redirected to the
Hugging Face mirror
via .lfsconfig, so dataset objects are fetched from Hugging
Face's CDN rather than from GitHub's shared (and limited) LFS bandwidth. This is
automatic — you do not need to configure anything.
Option 1: Clone the repository
git clone https://github.com/open-reaction-database/ord-data.git
With Git LFS installed… See the full description on the dataset page: https://huggingface.co/datasets/open-reaction-database/ord-data.smol-smoltalk-Interaction-SFT
Dataset Card for ReactiveAI/Smol-Smoltalk Interaction SFT
Derived from HuggingFaceTB/smol-smoltalk. Made for Interaction Supervised Fine-Tuning of Reactive Transformer
Proof-of-Concept models, especially RxT-Beta.
Dataset Details
Dataset Description
Reactive Transformers are processing only the single interactions in real-time and using Short-Term Memory to store information from previous interactions.
Before the model is able to use it's memory, it has to be… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/smol-smoltalk-Interaction-SFT.PDEBench_2D_diff-reactlegal:
owner: Takamoto, M et al. (https://darus.uni-stuttgart.de/dataset.xhtml?persistentId=doi:10.18419/darus-2986)
license: cc-by-4.0
data_production:
physics: 2D Diffusion-Reaction
type: simulation
script: Converted to PLAID format for standardized usage; no changes to data content.
num_samples:
train: 1000
storage_backend: hf_datasets
plaid:
version: 0.1.12
This dataset was generated with plaid, we refer to this documentation for additional details on how to extract data… See the full description on the dataset page: https://huggingface.co/datasets/Nionio/PDEBench_2D_diff-react.BioDEX-Reactions
Dataset Card for "BioDEX-Reactions"
More Information needed
hle-react
AggAgent ReAct Rollouts - HLE
Dataset Description
AggAgent is an agentic aggregation framework that scales long-horizon agents at test time by sampling multiple parallel rollouts from a base agent and then aggregating their evidence and solutions. This dataset card releases the ReAct base rollouts that AggAgent consumes, i.e. single-agent trajectories produced before any aggregation step.
Each rollout was generated by running a ReAct-style deep-research… See the full description on the dataset page: https://huggingface.co/datasets/yoonsanglee/hle-react.algebraic-stack-fixedBeta-Code
Reactive AI / Beta Code
Code-based pre-training corpus for RxT-Beta models, created from public & open datasets. Includes code in different programming languages.
Subsets are divided into short (< ~1024 tokens) and long (> ~1024 tokens) categories.
Original dataset
It's created from codeparrot datasets:
Python subsets from codeparrot/codeparrot-clean
other subsets from codeparrot/github-code-clean
RxQ-SMATfinepdfs-edu-betaReactionSmiles
Reaction SMILES Dataset
A collated dataset of 3.4M unique chemical reaction SMILES strings compiled from multiple public sources for use in pre-training and fine-tuning chemical language models.
Reaction SMILES (Simplified Molecular Input Line Entry System) extend the standard SMILES notation to represent complete chemical reactions. They encode reactants, reagents/catalysts, and products in a single text string using the > delimiter:
reactants>reagents>products
For example:… See the full description on the dataset page: https://huggingface.co/datasets/Derify/ReactionSmiles.RxQ-iSFTConversationalRetrieval-SMAT
ReactiveAI / ConverstationalRetrieval Dataset for Supervised Memory-Aware Training (SMAT)
Description in progress
beta-reasoninguniprot_reactions
Dataset Details
Dataset Description
Protein sequences and the reactions these can catalyze.
Curated by:
License: MIT
Dataset Sources
data source
Citation
BibTeX:
@article{10.1093/nar/gkac1052,
author = {The UniProt Consortium},
title = {UniProt - the Universal Protein Knowledgebase in 2023},
journal = {Nucleic Acids Research},
volume = {51},
number = {D1},
pages = {D523-D531},
year = {2022},
month = {11},
issn = {0305-1048},
doi =… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/uniprot_reactions.btc-news-forward-price-reactions
BTC News and Price Reactions
120,981 multilingual news events joined to weak semantic
annotations and BTC returns before and after each observation, from 1 minute
to 24 hours. The release is designed for event studies, classifier
bootstrapping, temporal evaluation, and research on how observed news aligns
with market movement.
Start here: licensed_news_price_reaction is the analysis-ready view with
attributed reusable titles/RSS summaries, weak labels, entry prices, and all
six… See the full description on the dataset page: https://huggingface.co/datasets/dymyt-ry/btc-news-forward-price-reactions.react-llama
The ReAct Llama Dataset
Dataset Summary
This dataset contains 3,538 correct ReAct trajectories generated using llama2-70b (Q5_K_M quant).
It follows the format used in the ReAct paper.ReAct trajectories were generated using a modified version of the hotpotqa.ipynb file from the ReAct repo.
The model was prompted in the following format (5-shot) to generate these traces:
Solve a question answering task with interleaving Thought, Action, Observation steps. Thought can… See the full description on the dataset page: https://huggingface.co/datasets/xz56/react-llama.Beta-Hybrid-SMAT
Reactive AI / Beta Hybrid SMAT
Multi-turn conversational dataset with hybrid reasoning for Supervised Memory Aware Training (SMAT) of Reactive Transformer MVP Beta models
Flame-Waterfall-React
Flame-Waterfall-React: A Structured Data Synthesis Dataset for Multimodal React Code Generation
Flame-Waterfall-React is a dataset synthesized using the Waterfall-Model-Based Synthesis method, Advancing Vision-Language Models in Front-End Development via Data Synthesis. This dataset is designed to train vision-language models (VLMs) for React code generation from UI design mockups and specifications.
The Waterfall synthesis approach mimics real-world software development by… See the full description on the dataset page: https://huggingface.co/datasets/Flame-Code-VLM/Flame-Waterfall-React.NVIDIA-Nemotron-IF-Chat-v2-rxNVIDIA-Nemotron-IF-Chat-v3-rx
README
omnimcp_nextjs_react_architect_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_nextjs_react_architect_teaser.ord-reactionsFlame-Additive-React
Flame-Additive-React: An Iterative Data Synthesis Dataset for Multi-modal React Code Generation
Flame-Additive-React is a dataset synthesized using the Additive Development Synthesis method, focusing on real-world React development patterns. This dataset ensures that training data remains grounded in realistic, incrementally enhanced code components.
Instead of generating synthetic data from scratch, this approach builds upon human-authored React components, progressively… See the full description on the dataset page: https://huggingface.co/datasets/Flame-Code-VLM/Flame-Additive-React.fineweb-2-pol-latest
ReactiveAI - FineWeb2 PL subset
This dataset is derived from polish subset of FineWeb2 by HuggingFace. Includes latest ~8.5M examples.
Original dataset description below
🥂 FineWeb2
A sparkling update with 1000s of languages
What is it?
This is the second iteration of the popular 🍷 FineWeb dataset, bringing high quality pretraining data to over 1000 🗣️ languages.
The 🥂 FineWeb2 dataset is fully reproducible, available under the permissive ODC-By… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/fineweb-2-pol-latest.deepsearchqa-react
AggAgent ReAct Rollouts - DeepSearchQA
Dataset Description
AggAgent is an agentic aggregation framework that scales long-horizon agents at test time by sampling multiple parallel rollouts from a base agent and then aggregating their evidence and solutions. This dataset card releases the ReAct base rollouts that AggAgent consumes, i.e. single-agent trajectories produced before any aggregation step.
Each rollout was generated by running a ReAct-style… See the full description on the dataset page: https://huggingface.co/datasets/yoonsanglee/deepsearchqa-react.context-conditioned-molecule-transfer-v10.3-skin-reaction-mixed-continuous-intern
Skin_Reaction context-conditioned molecule transfer V10.3
This release preserves its direct panels and appends training-only, parent-rendered V27 L1/L2/L3 transfer pairs after final-prompt cleanup. Query values remain hidden from prompts.
Train rows: 139,228
Validation rows: 18,167
Test rows: 18,319
emgena_nextjs_react19_fullstack_mcp_teaser
🚀 Full-Stack - Next.js 15 & React 19 Hydration & Server Action Guard (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Scenarios + Executable MCP Server)🏆 Get the Full Production Package & Commercial EULA on Gumroad:👉 Purchase Full Package on Gumroad🏷️ Use coupon code LAUNCH20 for €20 off at checkout!
🌟 Domain Overview & Features
SSR hydration mismatch resolution, Server Action secret leakage scanning, and async request waterfall… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_nextjs_react19_fullstack_mcp_teaser.
