Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ReactiveAI /smol-smoltalk-Interaction-SFT Dataset Card for ReactiveAI/Smol-Smoltalk Interaction SFT Derived from HuggingFaceTB/smol-smoltalk. Made for Interaction Supervised Fine-Tuning of Reactive Transformer Proof-of-Concept models, especially RxT-Beta. Dataset Details Dataset Description Reactive Transformers are processing only the single interactions in real-time and using Short-Term Memory to store information from previous interactions. Before the model is able to use it's memory, it has to be… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/smol-smoltalk-Interaction-SFT.texttext-generation1M<n<10M2 likes765 downloads1y agoHugging Face02yoonsanglee /hle-react AggAgent ReAct Rollouts - HLE Dataset Description AggAgent is an agentic aggregation framework that scales long-horizon agents at test time by sampling multiple parallel rollouts from a base agent and then aggregating their evidence and solutions. This dataset card releases the ReAct base rollouts that AggAgent consumes, i.e. single-agent trajectories produced before any aggregation step. Each rollout was generated by running a ReAct-style deep-research… See the full description on the dataset page: https://huggingface.co/datasets/yoonsanglee/hle-react.textquestion-answering1K<n<10K0 likes527 downloads5mo agoHugging Face03ReactiveAI /Beta-Code Reactive AI / Beta Code Code-based pre-training corpus for RxT-Beta models, created from public & open datasets. Includes code in different programming languages. Subsets are divided into short (< ~1024 tokens) and long (> ~1024 tokens) categories. Original dataset It's created from codeparrot datasets: Python subsets from codeparrot/codeparrot-clean other subsets from codeparrot/github-code-clean texttext-generation1M<n<10M0 likes340 downloads10mo agoHugging Face04Derify /ReactionSmiles Reaction SMILES Dataset A collated dataset of 3.4M unique chemical reaction SMILES strings compiled from multiple public sources for use in pre-training and fine-tuning chemical language models. Reaction SMILES (Simplified Molecular Input Line Entry System) extend the standard SMILES notation to represent complete chemical reactions. They encode reactants, reagents/catalysts, and products in a single text string using the > delimiter: reactants>reagents>products For example:… See the full description on the dataset page: https://huggingface.co/datasets/Derify/ReactionSmiles.texttext-generation1M<n<10M1 likes210 downloads9mo agoHugging Face05emgena /omnimcp_nextjs_react_architect_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_nextjs_react_architect_teaser.texttext-generationn<1K0 likes123 downloads24d agoHugging Face06ReactiveAI /fineweb-2-pol-latest ReactiveAI - FineWeb2 PL subset This dataset is derived from polish subset of FineWeb2 by HuggingFace. Includes latest ~8.5M examples. Original dataset description below 🥂 FineWeb2 A sparkling update with 1000s of languages What is it? This is the second iteration of the popular 🍷 FineWeb dataset, bringing high quality pretraining data to over 1000 🗣️ languages. The 🥂 FineWeb2 dataset is fully reproducible, available under the permissive ODC-By… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/fineweb-2-pol-latest.tabulartext-generation1M<n<10M0 likes105 downloads11mo agoHugging Face07yoonsanglee /deepsearchqa-react AggAgent ReAct Rollouts - DeepSearchQA Dataset Description AggAgent is an agentic aggregation framework that scales long-horizon agents at test time by sampling multiple parallel rollouts from a base agent and then aggregating their evidence and solutions. This dataset card releases the ReAct base rollouts that AggAgent consumes, i.e. single-agent trajectories produced before any aggregation step. Each rollout was generated by running a ReAct-style… See the full description on the dataset page: https://huggingface.co/datasets/yoonsanglee/deepsearchqa-react.textquestion-answering1K<n<10K0 likes91 downloads5mo agoHugging Face08ReactiveAI /RealStories-Micro-MRL Dataset Card for ReactiveAI/RealStories-Micro-MRL First synthetic Memory Reinforcement Learning dataset for Proof-of-Concept Reactive Transformer models. Dataset is divided into subsets, used in different Curriculum Stage of MRL training - each subset have different number of follow-up interactions, could use different strategy, and have train and validation splits. Subsets steps-1: ~2300 train (~4600 interactions) / ~340 validation (~680 interactions) - Single-Step… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/RealStories-Micro-MRL.textreinforcement-learning1K<n<10K0 likes81 downloads1y agoHugging Face09emgena /omnimcp_react_hydration_repair_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_react_hydration_repair_teaser.texttext-generationn<1K0 likes78 downloads24d agoHugging Face10ReactiveAI /TinyStories-MRL Dataset Card for ReactiveAI/TinyStories-MRL Synthetic Memory Reinforcement Learning dataset for Proof-of-Concept Reactive Transformer models. Dataset is divided into subsets, used in different Curriculum Stage of MRL training - each subset have different number of follow-up interactions, could use different strategy, and have train and validation splits. After first experiments with MRL, we decided to abandon single step and two steps stages. That's because with single step… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/TinyStories-MRL.textreinforcement-learning10K<n<100K0 likes72 downloads1y agoHugging Face11valentin-marquez /react-shadcn-codex React Shadcn Codex Dataset Description The React Shadcn Codex is a curated collection of over 3,000 React components that utilize shadcn, Framer Motion, and Lucide React. This dataset provides a valuable resource for developers looking to understand and implement modern React UI components with these popular libraries. Content The dataset includes: 3,000+ React components using shadcn UI Components with Framer Motion animations Usage examples of Lucide React… See the full description on the dataset page: https://huggingface.co/datasets/valentin-marquez/react-shadcn-codex.texttext-generation1K<n<10K13 likes67 downloads2y agoHugging Face12ReactiveAI /smol-smoltalk-mini-Interaction-SFT Dataset Card for ReactiveAI/Smol-Smoltalk-Mini Interaction SFT Derived from HuggingFaceTB/smol-smoltalk (used 25% of train & test splits). Made for Interaction Supervised Fine-Tuning of Reactive Transformer Proof-of-Concept models, especially RxT-Alpha-Mini (more info soon). Full version available in ReactiveAI/smol-smoltalk-Interaction-SFT Dataset Details Dataset Description Reactive Transformers are processing only the single interactions in… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/smol-smoltalk-mini-Interaction-SFT.textquestion-answering100K<n<1M0 likes64 downloads1y agoHugging Face13ReactiveAI /TinyStories-Interaction-SFT Dataset Card for ReactiveAI/TinyStories-Interaction-SFT Improved version of ReactiveAI/TinyStories-mini-Interaction-SFT - about 4x more rows, improved generation prompt and additional post-processing for more diverse dataset. Includes all examples from v1 Dataset and over 75k new ones, post-processed to include more random naming. Dataset Details Dataset Description Interaction Supervised Fine-Tuning Dataset based on roneneldan/TinyStories dataset, made… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/TinyStories-Interaction-SFT.texttext-generation100K<n<1M1 likes59 downloads1y agoHugging Face14ReactiveAI /TinyStories-mini-Interaction-SFT Dataset Card for ReactiveAI/TinyStories-mini-Interaction-SFT Interaction Supervised Fine-Tuning Dataset based on roneneldan/TinyStories dataset, made for Reactive Transformer second training stage Proof-of-Concept. Full version available in ReactiveAI/TinyStories-Interaction-SFT Dataset Details Dataset Description Curated by: Reactive AI Language(s) (NLP): English License: apache-2.0 Uses This dataset is made for Supervised… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/TinyStories-mini-Interaction-SFT.textquestion-answering10K<n<100K0 likes56 downloads1y agoHugging Face15ReactiveAI /AI-Knowledge-Chat-SMAT Dataset Card for ReactiveAI/AI-Knowledge-Chat-SMAT Conversational dataset for Supervised Memory Aware Training (SMAT) of Reactive Language Models (RxLM), containing dialogues with AI/Data Science knowledge. DOCS IN PROGRESS Dataset Details Dataset Description Curated by: Adam Filipek / Reactive AI Language(s) (NLP): English-only License: Apache-2.0 Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/AI-Knowledge-Chat-SMAT.textquestion-answering10K<n<100K2 likes52 downloads1y agoHugging Face16yoonsanglee /researchrubrics-react AggAgent ReAct Rollouts - ResearchRubrics Dataset Description AggAgent is an agentic aggregation framework that scales long-horizon agents at test time by sampling multiple parallel rollouts from a base agent and then aggregating their evidence and solutions. This dataset card releases the ReAct base rollouts that AggAgent consumes, i.e. single-agent trajectories produced before any aggregation step. Each rollout was generated by running a ReAct-style… See the full description on the dataset page: https://huggingface.co/datasets/yoonsanglee/researchrubrics-react.textquestion-answering1K<n<10K0 likes44 downloads5mo agoHugging Face17yoonsanglee /healthbench-react AggAgent ReAct Rollouts - HealthBench-Hard Dataset Description AggAgent is an agentic aggregation framework that scales long-horizon agents at test time by sampling multiple parallel rollouts from a base agent and then aggregating their evidence and solutions. This dataset card releases the ReAct base rollouts that AggAgent consumes, i.e. single-agent trajectories produced before any aggregation step. Each rollout was generated by running a ReAct-style… See the full description on the dataset page: https://huggingface.co/datasets/yoonsanglee/healthbench-react.textquestion-answering1K<n<10K0 likes43 downloads5mo agoHugging Face18ReactiveAI /passkey-retrieval ReactiveAI / passkey-retrieval (Interactions Format) Conversational (in RxLM Interactions Format) retrieval (Passkey / Needle In a Haystack type) dataset, filtered and transformed from grimulkan/passkey-retrieval Subsets to-4k - 3-step instruct examples with first (context) message with up to 4k tokens to-4k-reasoning - 3-step reasoning examples with first (context) query with up to 4k tokens and all the interaction (with reasoning) up to 8k tokens to-8k - 3-step… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/passkey-retrieval.texttext-retrieval1K<n<10K1 likes41 downloads6mo agoHugging Face19Azzindani /Open_Reaction_Data ORDerly: Styrene Mizoroki-Heck RAG-Ready Dataset This repository contains chemical reaction data formatted for Retrieval-Augmented Generation (RAG) systems. The data is a processed version of the ORDerly benchmark, specifically focusing on reaction conditions and forward/retro prediction tasks. Dataset Structure The data is split into 10,000-row Parquet chunks to prevent Out-of-Memory (OOM) errors during ingestion into vector databases. It includes: orderly_condition:… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/Open_Reaction_Data.tabulartext-generation1M<n<10M0 likes37 downloads7mo agoHugging Face20summykai /react-hook-form-sft-dataset React Hook Form-SFT: Instruction-Tuning Dataset for React Hook Form v7.68 Dataset: summykai/react-hook-form-sft-datasetFormat: ChatML-style messages + structured metadataLibrary: React Hook Form v7.68.0 (release date: 2025-12-03)License: MIT (react-hook-form) A research-grade supervised fine-tuning dataset that teaches models to write, debug, and explain React Hook Form patterns—covering both uncontrolled (register) and controlled (Controller) integrations, plus performance… See the full description on the dataset page: https://huggingface.co/datasets/summykai/react-hook-form-sft-dataset.texttext-generation1K<n<10K0 likes36 downloads10mo agoHugging Face21ReactiveAI /cosmopedia-100k-Interaction-SFT Dataset Card for ReactiveAI/cosmopedia-100k-Interaction-SFT HuggingFaceTB/cosmopedia-100k dataset transformed into Reactive AI Interaction SFT format for stateful real-time Reactive Language Models (RxLM). Dataset Details Dataset Description The subset of Cosmopedia with 100k examples for Interaction Supervised Fine-Tuning, made for Reactive Transformer models. We use it to train RxT-Beta generation. Curated by: Adam Filipek & Reactive AI Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/cosmopedia-100k-Interaction-SFT.textquestion-answering100K<n<1M1 likes35 downloads1y agoHugging Face22ReactorJet /coq-facts-props-proofs-gen0-v1 Dataset Name: Coq Facts, Propositions and Proofs Dataset Description The CoqFactsPropsProofs dataset aims to enhance Large Language Models' (LLMs) proficiency in interpreting and generating Coq code by providing a comprehensive collection of over 10,000 Coq source files. It encompasses a wide array of propositions, proofs, and definitions, enriched with metadata including source references and licensing information. This dataset is designed to facilitate the development… See the full description on the dataset page: https://huggingface.co/datasets/ReactorJet/coq-facts-props-proofs-gen0-v1.texttext-generation100K<n<1M0 likes34 downloads6mo agoHugging Face23chibifire /taskweft-fbd-react-train taskweft-fbd-react-train Intents and the IEC 61131-3 Function Block Diagrams that carry them out, as an EditScore-shaped corpus: one root row per intent, three candidates per row (rank1 the reference diagram, rank3 one that compiles and does the wrong thing, rank5 one the compiler refuses), and one score row per candidate from the compiler's reference scan on three constructed input traces per row. Every row is constructed from a template and a seed, so the labels are true by… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/taskweft-fbd-react-train.tabulartext-generation10K<n<100K0 likes30 downloads1mo agoHugging Face24ReactorJet /Nemotron-Math-Proofs-v1 Nemotron-Math-Proofs-v1 Paper: Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode SupervisionCode: https://github.com/NVIDIA/NeMo-SkillsDocumentation: Nemotron-MathProofs-v1 documentation Dataset Description: Nemotron-Math-Proofs-v1 is a large-scale mathematical reasoning dataset containing ~580k natural language proof problems, ~550k formalizations into theorem statements in Lean 4, and ~900k model-generated reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ReactorJet/Nemotron-Math-Proofs-v1.texttext-generation100K<n<1M0 likes26 downloads6mo agoHugging Face25onekq-ai /WebApp1K-React-Generations WebApp1K-React-Generations A comprehensive evaluation dataset containing React component generations from 34 state-of-the-art AI models on 1,000 web application scenarios. Dataset Description This dataset extends the original WebApp1K-React benchmark by including actual code generations from major AI models. Each row contains a web application scenario along with generated React components from 34 different models and their test results. Dataset Structure 1… See the full description on the dataset page: https://huggingface.co/datasets/onekq-ai/WebApp1K-React-Generations.texttext-generation1K<n<10K0 likes25 downloads1y agoHugging Face26u-10bei /dbbench_sft_dataset_react DBBench SFT Dataset (ReAct Format — AgentBench Compatible) Overview Synthetic SFT dataset for DBBench (AgentBench, ICLR 2024). All tables, data, and queries are independently generated to avoid test data leakage. Format ReAct text format matching the AgentBench DBBench evaluation protocol: [user] System prompt (Action: Operation / Action: Answer instructions) [agent] Ok. [user] Question + table name + column headers [agent] Thinking + Action: Operation +… See the full description on the dataset page: https://huggingface.co/datasets/u-10bei/dbbench_sft_dataset_react.texttext-generationn<1K0 likes23 downloads8mo agoHugging Face27onekq-ai /WebApp1K-Duo-React-Generations WebApp1K-Duo-React-Generations A comprehensive evaluation dataset containing React component generations from 32 state-of-the-art AI models on 1,000 paired web application scenarios. Dataset Description This dataset extends the original WebApp1K-Duo-React benchmark by including actual code generations from major AI models. Each row contains a paired web application scenario (combining two functionalities) along with generated React components from 32 different models and… See the full description on the dataset page: https://huggingface.co/datasets/onekq-ai/WebApp1K-Duo-React-Generations.texttext-generation1K<n<10K0 likes22 downloads1y agoHugging Face28u-10bei /dbbench_sft_dataset_react_v2 DBBench SFT Dataset (ReAct Format — AgentBench Compatible) Overview Synthetic SFT dataset for DBBench (AgentBench, ICLR 2024). All tables, data, and queries are independently generated to avoid test data leakage. Format ReAct text format matching the AgentBench DBBench evaluation protocol: [user] System prompt (Action: Operation / Action: Answer instructions) [agent] Ok. [user] Question + table name + column headers [agent] Thinking + Action: Operation +… See the full description on the dataset page: https://huggingface.co/datasets/u-10bei/dbbench_sft_dataset_react_v2.texttext-generationn<1K0 likes22 downloads8mo agoHugging Face29Salvor-Hardin /Search-Distill-5K-ReAct TraceMind-Search-5K TraceMind-Search-5K is a high-quality dataset featuring 5,000+ samples of search-augmented reasoning traces. It is specifically designed to train Small Language Models (SLMs) in the ReAct (Reasoning + Acting) pattern, focusing on precise information retrieval, noise filtering, and grounded synthesis. Dataset Highlights Total Samples: 4,997 entries in Parquet format. File: train.parquet (2.7 MB). Pattern: Strict adherence to the ReAct framework: User… See the full description on the dataset page: https://huggingface.co/datasets/Salvor-Hardin/Search-Distill-5K-ReAct.texttext-generation1K<n<10K0 likes22 downloads7mo agoHugging Face30u-10bei /dbbench_sft_dataset_react_v4 DBBench SFT Dataset (ReAct Format — AgentBench Compatible) Overview Synthetic SFT dataset for DBBench (AgentBench, ICLR 2024). All tables, data, and queries are independently generated to avoid test data leakage. Format ReAct text format matching the AgentBench DBBench evaluation protocol: [user] System prompt (Action: Operation / Action: Answer instructions) [agent] Ok. [user] Question + table name + column headers [agent] Thinking + Action: Operation +… See the full description on the dataset page: https://huggingface.co/datasets/u-10bei/dbbench_sft_dataset_react_v4.texttext-generation1K<n<10K0 likes20 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.