Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ReactiveAI /Beta-Hybrid-Interaction-SFTtext10M<n<100M0 likes16k downloads8mo agoHugging Face02ReactiveAI /Beta-Pre-Train-Corpus Reactive AI / Beta Pre-Train Corpus Pre-training corpus for RxT-Beta models, created from public & open datasets. Includes high-quality english and polish web crawl data, mathematic and scientific subsets, and code in different programming languages. 2k subsets are filtered for 1024-2048 tokens, except MegaMath Web Pro and GitHub Code subsets, that were filtered for 512-2048 tokens Subsets & original datasets FineWeb-Edu fineweb-edu-s100 (51.3M examples) - 50% of… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/Beta-Pre-Train-Corpus.tabular100M<n<1B0 likes15k downloads7mo agoHugging Face03yxma /React React — Multi-Task Tactile-Visual Manipulation Dense, contact-rich, synchronized multimodal interaction data collected from human hands holding handheld GelSight tactile sensors (no robot arm). Intended for tactile-visual dynamics / world-model learning. 133 min · 240 k frames @ 30 Hz · 3× RGB + 2× GelSight + OptiTrack · 2 tasks Format — LeRobot-style video release Each episode ships as 5 MP4 video streams (640×480, H.264) + a per-frame parquet of poses and… See the full description on the dataset page: https://huggingface.co/datasets/yxma/React.tabularrobotics1M<n<10M0 likes12k downloads13d agoHugging Face04open-reaction-database /ord-data ord-data Getting the Data The datasets live under data/ and are stored with Git LFS. LFS reads are redirected to the Hugging Face mirror via .lfsconfig, so dataset objects are fetched from Hugging Face's CDN rather than from GitHub's shared (and limited) LFS bandwidth. This is automatic — you do not need to configure anything. Option 1: Clone the repository git clone https://github.com/open-reaction-database/ord-data.git With Git LFS installed… See the full description on the dataset page: https://huggingface.co/datasets/open-reaction-database/ord-data.text1M<n<10M7 likes4.5k downloads1mo agoHugging Face05INV-WZQ /ReactiveGWM-Datasets ReactiveGWM-Datasets: Strategy-Aligned Rollouts for Reactive Game World Models 📚 Datasets-Introduction ReactiveGWM-Datasets is the strategy-aligned training corpus that powers ReactiveGWM, a game world model that decouples player control from NPC autonomy. To learn that decoupling, the model needs supervision that pairs each gameplay clip with both a per-frame action stream (what the player did) and a high-level NPC description (what the NPC tried to do, and under… See the full description on the dataset page: https://huggingface.co/datasets/INV-WZQ/ReactiveGWM-Datasets.textimage-to-video10K<n<100K8 likes2.9k downloads5mo agoHugging Face06Flame-Code-VLM /Flame-Waterfall-React Flame-Waterfall-React: A Structured Data Synthesis Dataset for Multimodal React Code Generation Flame-Waterfall-React is a dataset synthesized using the Waterfall-Model-Based Synthesis method, Advancing Vision-Language Models in Front-End Development via Data Synthesis. This dataset is designed to train vision-language models (VLMs) for React code generation from UI design mockups and specifications. The Waterfall synthesis approach mimics real-world software development by… See the full description on the dataset page: https://huggingface.co/datasets/Flame-Code-VLM/Flame-Waterfall-React.textimage-to-text100K<n<1M2 likes1k downloads1y agoHugging Face07Reubencf /webui-react-htmlcssjs-8740 WebUI React + HTML/CSS/JS 8,740 Curated export from ronantakizawa/webui containing every row where framework = react, plus 4,000 additional rows where framework = vanilla. Screenshots: 8,740 React / vanilla HTML-CSS-JS rows: 4,740 / 4,000 Unique sample IDs: 2,914 Train / validation / test: 7,501 / 456 / 783 Viewports: 2,914 desktop / 2,913 mobile / 2,913 tablet Images are stored as real image files and verified with Pillow. viewer.html is a self-contained, paginated local… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/webui-react-htmlcssjs-8740.imageimage-to-text1K<n<10K0 likes890 downloads2mo agoHugging Face08BioDEX /BioDEX-Reactions Dataset Card for "BioDEX-Reactions" More Information needed text10K<n<100K2 likes730 downloads3y agoHugging Face09openadmet /Octant_CYP_inhibition_reactivity_blog_release OpenADMET Octant CYP Inhibition & Reactivity Data release from the OpenADMET consortium, generated by Octant Bio. This dataset accompanies the blog post Building the OpenADMET Data Engine. Source code, assay protocols, and raw TSV files are on GitHub. Overview Cytochrome P450 (CYP) enzymes drive the oxidative metabolism of most drugs and are a primary cause of drug-drug interactions (DDIs). Despite their importance, public CYP datasets are sparse, noisy, and… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/Octant_CYP_inhibition_reactivity_blog_release.tabular10K<n<100K3 likes723 downloads2mo agoHugging Face10ReactiveAI /smol-smoltalk-Interaction-SFT Dataset Card for ReactiveAI/Smol-Smoltalk Interaction SFT Derived from HuggingFaceTB/smol-smoltalk. Made for Interaction Supervised Fine-Tuning of Reactive Transformer Proof-of-Concept models, especially RxT-Beta. Dataset Details Dataset Description Reactive Transformers are processing only the single interactions in real-time and using Short-Term Memory to store information from previous interactions. Before the model is able to use it's memory, it has to be… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/smol-smoltalk-Interaction-SFT.texttext-generation1M<n<10M2 likes690 downloads1y agoHugging Face11ReactiveAI /Beta-Code Reactive AI / Beta Code Code-based pre-training corpus for RxT-Beta models, created from public & open datasets. Includes code in different programming languages. Subsets are divided into short (< ~1024 tokens) and long (> ~1024 tokens) categories. Original dataset It's created from codeparrot datasets: Python subsets from codeparrot/codeparrot-clean other subsets from codeparrot/github-code-clean texttext-generation1M<n<10M0 likes505 downloads10mo agoHugging Face12ReactiveAI /algebraic-stack-fixedtext1M<n<10M0 likes454 downloads10mo agoHugging Face13yoonsanglee /hle-react AggAgent ReAct Rollouts - HLE Dataset Description AggAgent is an agentic aggregation framework that scales long-horizon agents at test time by sampling multiple parallel rollouts from a base agent and then aggregating their evidence and solutions. This dataset card releases the ReAct base rollouts that AggAgent consumes, i.e. single-agent trajectories produced before any aggregation step. Each rollout was generated by running a ReAct-style deep-research… See the full description on the dataset page: https://huggingface.co/datasets/yoonsanglee/hle-react.textquestion-answering1K<n<10K0 likes440 downloads5mo agoHugging Face14ReactiveAI /finepdfs-edu-betatabular10M<n<100M0 likes324 downloads11mo agoHugging Face15ReactiveAI /RxQ-SMATtext1M<n<10M0 likes324 downloads15d agoHugging Face16PureOne /pzhrd-programmable-zeno-holonomic-reaction-darkspace PZHRD — Programmable Zeno–Holonomic Reaction Darkspace Tangent-matched recovery, geometric reaction addressing, deferred-commit logical chemistry, and error-corrected matter construction Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiRelease: v1.0.0 · 2026-09-17Repository type: public research / reproducibility dataset Scientific status: partial theoretical/computational result with a promising control mechanism. This release does not demonstrate a universal… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/pzhrd-programmable-zeno-holonomic-reaction-darkspace.tabularothern<1K0 likes321 downloads19d agoHugging Face17cfahlgren1 /react-code-instructions React Code Instructions Popular Queries Number of instructions by Model Unnested Messages Instructions Added Per Day Dataset of Claude Artifact esque React Apps generated by Llama 3.1 70B, Llama 3.1 405B, and Deepseek Chat V3. Examples Virtual Fitness Trainer Website LinkedIn Clone iPhone Calculator Chipotle Waitlist Apple Store text10K<n<100K158 likes308 downloads2y agoHugging Face18Flame-Code-VLM /Flame-Additive-React Flame-Additive-React: An Iterative Data Synthesis Dataset for Multi-modal React Code Generation Flame-Additive-React is a dataset synthesized using the Additive Development Synthesis method, focusing on real-world React development patterns. This dataset ensures that training data remains grounded in realistic, incrementally enhanced code components. Instead of generating synthetic data from scratch, this approach builds upon human-authored React components, progressively… See the full description on the dataset page: https://huggingface.co/datasets/Flame-Code-VLM/Flame-Additive-React.textimage-to-text10K<n<100K3 likes232 downloads1y agoHugging Face19ReactiveAI /RxQ-iSFTtext1M<n<10M0 likes230 downloads3mo agoHugging Face20scikit-fingerprints /TDC_skin_reaction TDC Skin Reaction Skin Reaction dataset dataset [1], part of TDC [2] benchmark. It is intended to be used through scikit-fingerprints library. The task is to predict whether the drug can cause immune reaction that leads to skin sensitization. Characteristic Description Tasks 1 Task type classification Total samples 404 Recommended split scaffold Recommended metricAUROC References [1] Alves, Vinicius M., et al. "Predicting chemically-induced skin… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/TDC_skin_reaction.texttabular-classificationn<1K0 likes224 downloads1y agoHugging Face21crystantine /sd-reactor-node ReActor Node 0.1.1b for ComfyUI The Fast and Simple "roop-like" Face Swap Extension Node for ComfyUI, based on ReActor (ex Roop-GE) SD-WebUI Face Swap Extension This Node goes without NSFW filter (uncensored, use it on your own responsibility) Disclaimer | Installation | Usage | Troubleshooting | Updating Disclaimer This software is meant to be a productive contribution to the rapidly growing AI-generated media industry. It will help artists with tasks… See the full description on the dataset page: https://huggingface.co/datasets/crystantine/sd-reactor-node.imagen<1K7 likes218 downloads3y agoHugging Face22jablonkagroup /uniprot_reactions Dataset Details Dataset Description Protein sequences and the reactions these can catalyze. Curated by: License: MIT Dataset Sources data source Citation BibTeX: @article{10.1093/nar/gkac1052, author = {The UniProt Consortium}, title = {UniProt - the Universal Protein Knowledgebase in 2023}, journal = {Nucleic Acids Research}, volume = {51}, number = {D1}, pages = {D523-D531}, year = {2022}, month = {11}, issn = {0305-1048}, doi =… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/uniprot_reactions.text1M<n<10M0 likes207 downloads1y agoHugging Face23Derify /ReactionSmiles Reaction SMILES Dataset A collated dataset of 3.4M unique chemical reaction SMILES strings compiled from multiple public sources for use in pre-training and fine-tuning chemical language models. Reaction SMILES (Simplified Molecular Input Line Entry System) extend the standard SMILES notation to represent complete chemical reactions. They encode reactants, reagents/catalysts, and products in a single text string using the > delimiter: reactants>reagents>products For example:… See the full description on the dataset page: https://huggingface.co/datasets/Derify/ReactionSmiles.texttext-generation1M<n<10M1 likes200 downloads9mo agoHugging Face24ReactiveAI /ConversationalRetrieval-SMAT ReactiveAI / ConverstationalRetrieval Dataset for Supervised Memory-Aware Training (SMAT) Description in progress text100K<n<1M0 likes197 downloads5mo agoHugging Face25ReactiveAI /beta-reasoningtext100K<n<1M0 likes181 downloads3mo agoHugging Face26ReactiveAI /Beta-Hybrid-SMAT Reactive AI / Beta Hybrid SMAT Multi-turn conversational dataset with hybrid reasoning for Supervised Memory Aware Training (SMAT) of Reactive Transformer MVP Beta models text100K<n<1M0 likes171 downloads6mo agoHugging Face27ReactiveAI /NVIDIA-Nemotron-IF-Chat-v3-rx README text100K<n<1M0 likes163 downloads3mo agoHugging Face28iheallab /ICU-REACT ICU-REACT ICU-REACT is a clinician-supervised dataset for clinical reasoning and information retrieval in the intensive care unit (ICU), developed for fine-tuning and benchmarking large language models (LLMs). ICU-REACT was constructed using a clinician-in-the-loop annotation framework designed to capture how clinicians identify relevant patient information and integrate it into diagnostic and treatment decisions. The dataset includes a clinician-refined seed training set, a… See the full description on the dataset page: https://huggingface.co/datasets/iheallab/ICU-REACT.textquestion-answering10K<n<100K1 likes154 downloads1mo agoHugging Face29smitathkr1 /ord-reactionstext1M<n<10M0 likes149 downloads11mo agoHugging Face30chibifire /taskweft-fbd-react-train taskweft-fbd-react-train Intents and the IEC 61131-3 Function Block Diagrams that carry them out, as an EditScore-shaped corpus: one root row per intent, three candidates per row (rank1 the reference diagram, rank3 one that compiles and does the wrong thing, rank5 one the compiler refuses), and one score row per candidate from the compiler's reference scan on three constructed input traces per row. Every row is constructed from a template and a seed, so the labels are true by… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/taskweft-fbd-react-train.tabulartext-generation10K<n<100K0 likes147 downloads28d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.