Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ReactiveAI /Beta-Pre-Train-Corpus Reactive AI / Beta Pre-Train Corpus Pre-training corpus for RxT-Beta models, created from public & open datasets. Includes high-quality english and polish web crawl data, mathematic and scientific subsets, and code in different programming languages. 2k subsets are filtered for 1024-2048 tokens, except MegaMath Web Pro and GitHub Code subsets, that were filtered for 512-2048 tokens Subsets & original datasets FineWeb-Edu fineweb-edu-s100 (51.3M examples) - 50% of… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/Beta-Pre-Train-Corpus.tabular100M<n<1B0 likes13k downloads8mo agoHugging Face02ReactiveAI /Beta-Hybrid-Interaction-SFTtext10M<n<100M0 likes13k downloads8mo agoHugging Face03yxma /React React — Multi-Task Tactile-Visual Manipulation Dense, contact-rich, synchronized multimodal interaction data collected from human hands holding handheld GelSight tactile sensors (no robot arm). Intended for tactile-visual dynamics / world-model learning. 133 min · 240 k frames @ 30 Hz · 3× RGB + 2× GelSight + OptiTrack · 2 tasks Format — LeRobot-style video release Each episode ships as 5 MP4 video streams (640×480, H.264) + a per-frame parquet of poses and… See the full description on the dataset page: https://huggingface.co/datasets/yxma/React.tabularrobotics1M<n<10M0 likes12k downloads17d agoHugging Face04open-reaction-database /ord-data ord-data Getting the Data The datasets live under data/ and are stored with Git LFS. LFS reads are redirected to the Hugging Face mirror via .lfsconfig, so dataset objects are fetched from Hugging Face's CDN rather than from GitHub's shared (and limited) LFS bandwidth. This is automatic — you do not need to configure anything. Option 1: Clone the repository git clone https://github.com/open-reaction-database/ord-data.git With Git LFS installed… See the full description on the dataset page: https://huggingface.co/datasets/open-reaction-database/ord-data.text1M<n<10M8 likes4.8k downloads1mo agoHugging Face05ReactiveAI /smol-smoltalk-Interaction-SFT Dataset Card for ReactiveAI/Smol-Smoltalk Interaction SFT Derived from HuggingFaceTB/smol-smoltalk. Made for Interaction Supervised Fine-Tuning of Reactive Transformer Proof-of-Concept models, especially RxT-Beta. Dataset Details Dataset Description Reactive Transformers are processing only the single interactions in real-time and using Short-Term Memory to store information from previous interactions. Before the model is able to use it's memory, it has to be… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/smol-smoltalk-Interaction-SFT.texttext-generation1M<n<10M2 likes765 downloads1y agoHugging Face06Nionio /PDEBench_2D_diff-reactlegal: owner: Takamoto, M et al. (https://darus.uni-stuttgart.de/dataset.xhtml?persistentId=doi:10.18419/darus-2986) license: cc-by-4.0 data_production: physics: 2D Diffusion-Reaction type: simulation script: Converted to PLAID format for standardized usage; no changes to data content. num_samples: train: 1000 storage_backend: hf_datasets plaid: version: 0.1.12 This dataset was generated with plaid, we refer to this documentation for additional details on how to extract data… See the full description on the dataset page: https://huggingface.co/datasets/Nionio/PDEBench_2D_diff-react.timeseriesgraph-ml1K<n<10K0 likes684 downloads4mo agoHugging Face07BioDEX /BioDEX-Reactions Dataset Card for "BioDEX-Reactions" More Information needed text10K<n<100K2 likes648 downloads3y agoHugging Face08yoonsanglee /hle-react AggAgent ReAct Rollouts - HLE Dataset Description AggAgent is an agentic aggregation framework that scales long-horizon agents at test time by sampling multiple parallel rollouts from a base agent and then aggregating their evidence and solutions. This dataset card releases the ReAct base rollouts that AggAgent consumes, i.e. single-agent trajectories produced before any aggregation step. Each rollout was generated by running a ReAct-style deep-research… See the full description on the dataset page: https://huggingface.co/datasets/yoonsanglee/hle-react.textquestion-answering1K<n<10K0 likes527 downloads5mo agoHugging Face09ReactiveAI /algebraic-stack-fixedtext1M<n<10M0 likes374 downloads10mo agoHugging Face10ReactiveAI /Beta-Code Reactive AI / Beta Code Code-based pre-training corpus for RxT-Beta models, created from public & open datasets. Includes code in different programming languages. Subsets are divided into short (< ~1024 tokens) and long (> ~1024 tokens) categories. Original dataset It's created from codeparrot datasets: Python subsets from codeparrot/codeparrot-clean other subsets from codeparrot/github-code-clean texttext-generation1M<n<10M0 likes340 downloads10mo agoHugging Face11ReactiveAI /RxQ-SMATtext1M<n<10M0 likes305 downloads19d agoHugging Face12ReactiveAI /finepdfs-edu-betatabular10M<n<100M0 likes291 downloads11mo agoHugging Face13Derify /ReactionSmiles Reaction SMILES Dataset A collated dataset of 3.4M unique chemical reaction SMILES strings compiled from multiple public sources for use in pre-training and fine-tuning chemical language models. Reaction SMILES (Simplified Molecular Input Line Entry System) extend the standard SMILES notation to represent complete chemical reactions. They encode reactants, reagents/catalysts, and products in a single text string using the > delimiter: reactants>reagents>products For example:… See the full description on the dataset page: https://huggingface.co/datasets/Derify/ReactionSmiles.texttext-generation1M<n<10M1 likes210 downloads9mo agoHugging Face14ReactiveAI /RxQ-iSFTtext1M<n<10M0 likes197 downloads3mo agoHugging Face15ReactiveAI /ConversationalRetrieval-SMAT ReactiveAI / ConverstationalRetrieval Dataset for Supervised Memory-Aware Training (SMAT) Description in progress text100K<n<1M0 likes184 downloads5mo agoHugging Face16ReactiveAI /beta-reasoningtext100K<n<1M0 likes182 downloads3mo agoHugging Face17jablonkagroup /uniprot_reactions Dataset Details Dataset Description Protein sequences and the reactions these can catalyze. Curated by: License: MIT Dataset Sources data source Citation BibTeX: @article{10.1093/nar/gkac1052, author = {The UniProt Consortium}, title = {UniProt - the Universal Protein Knowledgebase in 2023}, journal = {Nucleic Acids Research}, volume = {51}, number = {D1}, pages = {D523-D531}, year = {2022}, month = {11}, issn = {0305-1048}, doi =… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/uniprot_reactions.text1M<n<10M0 likes180 downloads1y agoHugging Face18dymyt-ry /btc-news-forward-price-reactions BTC News and Price Reactions 120,981 multilingual news events joined to weak semantic annotations and BTC returns before and after each observation, from 1 minute to 24 hours. The release is designed for event studies, classifier bootstrapping, temporal evaluation, and research on how observed news aligns with market movement. Start here: licensed_news_price_reaction is the analysis-ready view with attributed reusable titles/RSS summaries, weak labels, entry prices, and all six… See the full description on the dataset page: https://huggingface.co/datasets/dymyt-ry/btc-news-forward-price-reactions.tabulartext-classification1M<n<10M0 likes158 downloads18d agoHugging Face19xz56 /react-llama The ReAct Llama Dataset Dataset Summary This dataset contains 3,538 correct ReAct trajectories generated using llama2-70b (Q5_K_M quant). It follows the format used in the ReAct paper.ReAct trajectories were generated using a modified version of the hotpotqa.ipynb file from the ReAct repo. The model was prompted in the following format (5-shot) to generate these traces: Solve a question answering task with interleaving Thought, Action, Observation steps. Thought can… See the full description on the dataset page: https://huggingface.co/datasets/xz56/react-llama.text1K<n<10K55 likes145 downloads3y agoHugging Face20ReactiveAI /Beta-Hybrid-SMAT Reactive AI / Beta Hybrid SMAT Multi-turn conversational dataset with hybrid reasoning for Supervised Memory Aware Training (SMAT) of Reactive Transformer MVP Beta models text100K<n<1M0 likes139 downloads6mo agoHugging Face21Flame-Code-VLM /Flame-Waterfall-React Flame-Waterfall-React: A Structured Data Synthesis Dataset for Multimodal React Code Generation Flame-Waterfall-React is a dataset synthesized using the Waterfall-Model-Based Synthesis method, Advancing Vision-Language Models in Front-End Development via Data Synthesis. This dataset is designed to train vision-language models (VLMs) for React code generation from UI design mockups and specifications. The Waterfall synthesis approach mimics real-world software development by… See the full description on the dataset page: https://huggingface.co/datasets/Flame-Code-VLM/Flame-Waterfall-React.textimage-to-text100K<n<1M2 likes135 downloads1y agoHugging Face22ReactiveAI /NVIDIA-Nemotron-IF-Chat-v2-rxtext1M<n<10M0 likes135 downloads3mo agoHugging Face23ReactiveAI /NVIDIA-Nemotron-IF-Chat-v3-rx README text100K<n<1M0 likes132 downloads3mo agoHugging Face24emgena /omnimcp_nextjs_react_architect_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_nextjs_react_architect_teaser.texttext-generationn<1K0 likes123 downloads24d agoHugging Face25smitathkr1 /ord-reactionstext1M<n<10M0 likes118 downloads11mo agoHugging Face26Flame-Code-VLM /Flame-Additive-React Flame-Additive-React: An Iterative Data Synthesis Dataset for Multi-modal React Code Generation Flame-Additive-React is a dataset synthesized using the Additive Development Synthesis method, focusing on real-world React development patterns. This dataset ensures that training data remains grounded in realistic, incrementally enhanced code components. Instead of generating synthetic data from scratch, this approach builds upon human-authored React components, progressively… See the full description on the dataset page: https://huggingface.co/datasets/Flame-Code-VLM/Flame-Additive-React.textimage-to-text10K<n<100K3 likes106 downloads1y agoHugging Face27ReactiveAI /fineweb-2-pol-latest ReactiveAI - FineWeb2 PL subset This dataset is derived from polish subset of FineWeb2 by HuggingFace. Includes latest ~8.5M examples. Original dataset description below 🥂 FineWeb2 A sparkling update with 1000s of languages What is it? This is the second iteration of the popular 🍷 FineWeb dataset, bringing high quality pretraining data to over 1000 🗣️ languages. The 🥂 FineWeb2 dataset is fully reproducible, available under the permissive ODC-By… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/fineweb-2-pol-latest.tabulartext-generation1M<n<10M0 likes105 downloads11mo agoHugging Face28yoonsanglee /deepsearchqa-react AggAgent ReAct Rollouts - DeepSearchQA Dataset Description AggAgent is an agentic aggregation framework that scales long-horizon agents at test time by sampling multiple parallel rollouts from a base agent and then aggregating their evidence and solutions. This dataset card releases the ReAct base rollouts that AggAgent consumes, i.e. single-agent trajectories produced before any aggregation step. Each rollout was generated by running a ReAct-style… See the full description on the dataset page: https://huggingface.co/datasets/yoonsanglee/deepsearchqa-react.textquestion-answering1K<n<10K0 likes91 downloads5mo agoHugging Face29jiosephlee /context-conditioned-molecule-transfer-v10.3-skin-reaction-mixed-continuous-intern Skin_Reaction context-conditioned molecule transfer V10.3 This release preserves its direct panels and appends training-only, parent-rendered V27 L1/L2/L3 transfer pairs after final-prompt cleanup. Query values remain hidden from prompts. Train rows: 139,228 Validation rows: 18,167 Test rows: 18,319 tabular100K<n<1M0 likes87 downloads20d agoHugging Face30emgena /emgena_nextjs_react19_fullstack_mcp_teaser 🚀 Full-Stack - Next.js 15 & React 19 Hydration & Server Action Guard (Evaluation Teaser) ⚡ Official Free Evaluation Teaser (50 Verified Scenarios + Executable MCP Server)🏆 Get the Full Production Package & Commercial EULA on Gumroad:👉 Purchase Full Package on Gumroad🏷️ Use coupon code LAUNCH20 for €20 off at checkout! 🌟 Domain Overview & Features SSR hydration mismatch resolution, Server Action secret leakage scanning, and async request waterfall… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_nextjs_react19_fullstack_mcp_teaser.textn<1K0 likes84 downloads10d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.