Team Ai
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mdonigian /full-structured-instruction-sft-dataset Full Structured + Instruction SFT Corpus Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data. Dataset repo mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11 Included files train_full_sft.jsonl: full merged and shuffled SFT dataset source_glaive.jsonl: processed Glaive subset source_hermes.jsonl: processed Hermes subset source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.texttext-generation10K<n<100K0 likes160 downloads7mo agoHugging Face02Z-Edgar /Agent-IPI-Structured-Interaction-Datasets Dataset Card for Indirect Prompt Injection in Agent Structured Interaction Datasets Dataset Summary This dataset contains 470,000 QA pairs designed to study indirect prompt injection in agent-structured interactions. It is split into a training set (80%) and a test set (20%). The dataset is evenly divided into 50% clean-clean QA pairs (no prompt injection) and 50% clean-injected QA pairs (containing prompt injection). The task is to detect and remove prompt injection… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/Agent-IPI-Structured-Interaction-Datasets.text100K<n<1M1 likes90 downloads10mo agoHugging Face03u-10bei /structured_data_with_cot_dataset_512_v4 structured_data_with_cot_dataset このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。 データセットの概要 messages: OpenAIチャット形式 (system, user, assistant) metadata: format, complexity, schema, estimated_tokens サポートされるデータ形式 JSON, XML, YAML, TOML, CSV 生成方法 Fakerライブラリを使用し、Pythonスクリプトで生成。検証用・テスト用に分割済み。 text1K<n<10K1 likes56 downloads9mo agoHugging Face04u-10bei /structured_data_with_cot_dataset_512_v5 structured_data_with_cot_dataset このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。 データセットの概要 messages: OpenAIチャット形式 (system, user, assistant) metadata: format, complexity, schema, estimated_tokens サポートされるデータ形式 JSON, XML, YAML, TOML, CSV 生成方法 Fakerライブラリを使用し、Pythonスクリプトで生成。検証用・テスト用に分割済み。 v5アップデート:ランダムなスキーマ構造の生成と、最小化(minified)/ソート(sorted)の制約を追加。 text1K<n<10K3 likes31 downloads9mo agoHugging Face05ToshiyukiNH /structured_data_with_cot_dataset_512_v2_filtered_1structured_data_with_cot_dataset_512_v2_filtered_1 This repository provides a dataset for training models in terms of strutured outputs. The dataset is a part of u-10bei/structured_data_with_cot_dataset_512_v2. Usage from datasets import load_dataset # From local data dataset = load_dataset( 'json', data_files="ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_1", split='train' ) print(dataset[0]) How to generate this dataset from the base one from… See the full description on the dataset page: https://huggingface.co/datasets/ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_1.textn<1K0 likes28 downloads7mo agoHugging Face06u-10bei /structured_data_with_cot_dataset_512_v3 structured_data_with_cot_dataset このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。 データセットの概要 messages: OpenAIチャット形式 (system, user, assistant) metadata: format, complexity, schema, estimated_tokens サポートされるデータ形式 JSON, XML, YAML, TOML, CSV 生成方法 Fakerライブラリを使用し、Pythonスクリプトで生成。検証用・テスト用に分割済み。 text1K<n<10K0 likes27 downloads9mo agoHugging Face07Nexdata-AI /1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample Description Korean Test Questions Structured Analysis Processing Data, around 1.5 million questions, contains question types, questions, answers, explanations, etc..For subjects, include [Primary School] Korean, Mathematics, English, Social Studies, Science; [Middle School] Korean, English, Mathematics, Science, Social Studies; [High School] Korean, English, Mathematics, Physics, Chemistry, Biology, History, Geography; question Types indlude single-choice question, fill-in… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.textn<1K0 likes23 downloads2mo agoHugging Face08ToshiyukiNH /structured_data_with_cot_dataset_512_v2_filtered_3structured_data_with_cot_dataset_512_v2_filtered_3 This repository provides a dataset for training models in terms of strutured outputs. The dataset is a part of u-10bei/structured_data_with_cot_dataset_512_v2. Usage from datasets import load_dataset # From local data dataset = load_dataset( 'json', data_files="ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_3", split='train' ) print(dataset[0]) How to generate this dataset from the base one from… See the full description on the dataset page: https://huggingface.co/datasets/ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_3.textn<1K0 likes17 downloads7mo agoHugging Face09ToshiyukiNH /structured_data_with_cot_dataset_512_v2_filtered_4structured_data_with_cot_dataset_512_v2_filtered_4 This repository provides a dataset for training models in terms of strutured outputs. The dataset is a part of u-10bei/structured_data_with_cot_dataset_512_v2. Usage from datasets import load_dataset # From local data dataset = load_dataset( 'json', data_files="ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_4", split='train' ) print(dataset[0]) How to generate this dataset from the base one from… See the full description on the dataset page: https://huggingface.co/datasets/ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_4.text1K<n<10K0 likes15 downloads7mo agoHugging Face10ariefansclub /humanoid-domestic-task-structured-dataset-v2 Humanoid Domestic Task Structured Dataset Overview This dataset contains structured human instructions for basic household assistance scenarios. It is designed to help humanoid agents interpret natural language commands and convert them into clear executable task representations. The dataset focuses on simple real-world domestic tasks that reduce human workload and improve everyday living environments. Key Features Natural human-written instructions Structured… See the full description on the dataset page: https://huggingface.co/datasets/ariefansclub/humanoid-domestic-task-structured-dataset-v2.textn<1K0 likes13 downloads8mo agoHugging Face11ToshiyukiNH /structured_data_with_cot_dataset_512_v2_filtered_2structured_data_with_cot_dataset_512_v2_filtered_2 This repository provides a dataset for training models in terms of strutured outputs. The dataset is a part of u-10bei/structured_data_with_cot_dataset_512_v2. Usage from datasets import load_dataset # From local data dataset = load_dataset( 'json', data_files="ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_2", split='train' ) print(dataset[0]) How to generate this dataset from the base one from… See the full description on the dataset page: https://huggingface.co/datasets/ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_2.textn<1K0 likes13 downloads7mo agoHugging Face12OsakanaTeishoku /structured_data_with_cot_dataset_512_v2_dpotext1K<n<10K0 likes11 downloads9mo agoHugging Face13Nexdata-kr /1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample Description 한국어 시험 문제 구조화 분석·가공 데이터로, 약 150만 개의 시험 문제를 포함하고 있습니다. 문제 유형, 문제, 정답, 해설 등의 정보를 포함하며, 과목은 [초등학교] 국어, 수학, 영어, 사회, 과학; [중학교] 국어, 영어, 수학, 과학, 사회; [고등학교] 국어, 영어, 수학, 물리, 화학, 생물, 역사, 지리로 구성되어 있습니다. 문제 유형에는 객관식, 빈칸 채우기, 참·거짓 문제, 단답형 문제 등이 포함됩니다. 본 데이터셋은 대규모 교과 지식 강화 및 학습 데이터 구축 등의 작업에 활용할 수 있습니다. 자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/llm/1634?source=Hf.kr Specifications Data content 한국어 K12 시험 문제 Amount 약 150만 개의… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.textn<1K0 likes11 downloads1mo agoHugging Face14dheerajpabolu /Indian_Laws_Structured_Legal_Dataset 📚 Indian Legal Acts Dataset (Structured Sections) 🧾 Overview This dataset provides structured, machine-readable legal text from major Indian statutes, including: Bharatiya Nyaya Sanhita, 2023 (BNS) Code of Criminal Procedure, 1973 (CrPC) Code of Civil Procedure, 1908 (CPC) Indian Evidence Act, 1872 (IEA) Negotiable Instruments Act, 1881 (NIA) Motor Vehicles Act, 1988 (MVA) Indian Divorce Act, 1869 (IDA) Each entry represents a section or chunk of a section… See the full description on the dataset page: https://huggingface.co/datasets/dheerajpabolu/Indian_Laws_Structured_Legal_Dataset.tabular1K<n<10K0 likes9 downloads6mo agoHugging Face15sadnjasdkn /JSON-STRUCTURED-DATA-FOR-SYMPTOMS-SFT_DPO-SUPPORTEDThis dataset is high-consistency instruction tuning dataset. converting messy, subjective human health-style text → structured, non-diagnostic extraction format 1.Literal extraction discipline 2.Source separation logic - very strong schema grounding training if DPO 3.Anti-inference constraint 1.High ambiguity coverage 2.Contradiction handling included 3.Minimization bias detection This dataset is a STRICT schema regulation. Does well at: strict extraction preserving uncertainty words… See the full description on the dataset page: https://huggingface.co/datasets/sadnjasdkn/JSON-STRUCTURED-DATA-FOR-SYMPTOMS-SFT_DPO-SUPPORTED.text1K<n<10K0 likes7 downloads4mo agoHugging Face16OsakanaTeishoku /structured_data_with_cot_dataset_512_v2_dpo_before_processingtext1K<n<10K0 likes5 downloads9mo agoHugging Face17ShubhangiMishra /finetune-Sample-Structured-Datasettextn<1K0 likes2 downloads2y agoHugging Face18angiedev /structured_datasettextn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.