datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AgentDoG1.0-Training-Data
AgentDoG1.0 Training Data
[💻 GitHub] | [📊 ATBench Dataset] | [📄 ATBench Paper] | [📄 AgentDoG Paper] | [🤗 Collection]
AgentDoG1.0 Training Data releases supervised instruction-tuning data for trajectory-level AI-agent safety modeling. It is paired with the AgentDoG and ATBench line of work: ATBench is the benchmark release, while this repository contains training-oriented data for binary safety classification and fine-grained taxonomy diagnosis.
Introduction… See the full description on the dataset page: https://huggingface.co/datasets/AI45Research/AgentDoG1.0-Training-Data.Agent-G2-ALFWorld-Webshop-sft-data
Agent-G2 SFT Data
Agent-G2 SFT Data contains reasoning and action trajectories for supervised
fine-tuning (SFT) in the Agent-G2 project.
Associated paper: Agent-G2: Gaussian Guidance for Agentic Reinforcement
Learning — accepted to the EMNLP 2026 Main Conference.
The dataset covers two interactive agent environments:
WebShop: agents search for products, select options, and complete
purchases according to user requirements.
ALFWorld: agents interact with household environments… See the full description on the dataset page: https://huggingface.co/datasets/xiamoent/Agent-G2-ALFWorld-Webshop-sft-data.sft-dataproper-agents-data
ProPer Agents — data
Data for ProPer Agents: Proactivity Driven Personalized Agents for Advancing
Knowledge Gap Navigation (ACL 2026).
Paper ·
Adapters
Three domains: code, medical, pwab (product recommendation).
Layout
{domain}/
raw/train.jsonl source examples
raw/test.jsonl
raw/{domain}_rga_{train,test}.jsonl RGA SFT data (Alpaca format)
raw/{domain}_dga_{train,test}.jsonl DGA SFT data (Alpaca format)… See the full description on the dataset page: https://huggingface.co/datasets/itsgupta/proper-agents-data.Pretrain-Taiwan-DentistKnowledge-zhTW-290KLaplaceAI 繁中領域知識資料集計畫
利用我在爬蟲自動化與資料後處理上的專業,針對不同大小的領域知識資料集進行建立與維護。
在 LaplaceAI 的 huggingface 頁面,你可以找到許多不同領域的資料集。
這項 datasets 是由 LaplaceAI 整理維護的牙科相關知識。
agentdog-lite-qwen35-08b-base-training-data-suite
AgentDoG-Lite Qwen3.5-0.8B 基座训练数据套件
本数据集用于 AgentDoG-Lite Summer Camp 轨迹级 Agent 安全诊断任务,目标是训练模型判断完整 agent trajectory 是否安全。
核心判断标准不是风险词匹配,而是:
Agent 是否实际执行了 unsafe action。
即:
风险出现 != unsafe
风险被执行 == unsafe
相关模型
Full-SFT 完整权重:https://huggingface.co/hhhggfdd/doc-was-wrong-because-training-started-from-qwen3.5-0.8b-base-not-agentdog-full-sft
LoRA adapter:https://huggingface.co/hhhggfdd/doc-was-wrong-because-training-started-from-qwen3.5-0.8b-base-not-agentdog-lora… See the full description on the dataset page: https://huggingface.co/datasets/hhhggfdd/agentdog-lite-qwen35-08b-base-training-data-suite.ptdbench-data-format-task-agent-loop-022-llama-dapo-math-dataset
PTDBench dataset snapshot: task_agent_loop_022-llama-dapo-math
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: data_format
Source evaluation metric: val-core/math_dapo/reward/mean@1
Provenance: Processed from BytedTsinghua-SIA/DAPO-Math-17k; task-specific bytes are pinned.
License: Apache-2.0
The artifact manifest records every hydrated runtime… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-data-format-task-agent-loop-022-llama-dapo-math-dataset.TCNNet-SFT-NetCom-zhTW-1.1M
[TCNNet] A Traditional Chinese Networking and Communication Instruction Fine-Tuning Dataset (zh-TW)
A large-scale supervised fine-tuning (SFT) dataset created specifically for TCNNet-9B, a Chinese language model specialized in networking and communications domains. The dataset contains question-answer pairs generated from various networking, cybersecurity, and tech review articles written in Traditional Chinese.
Dataset Description
Dataset Summary
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/DataAgent/TCNNet-SFT-NetCom-zhTW-1.1M.TCNNet-Pretrain-NetCom-zhTW-3.7M
[TCNNet] A Large-scale Traditional Chinese Networking and Communication Continuous Pretraining Dataset (zh-TW)
A specialized domain knowledge dataset created for continuous pretraining of TCNNet-9B, a Chinese language model based on Yi-9B and specialized in networking and communications domains. The dataset contains articles from various networking, cybersecurity, and tech review sources written in Traditional Chinese.
Dataset Description
Dataset Summary
This… See the full description on the dataset page: https://huggingface.co/datasets/DataAgent/TCNNet-Pretrain-NetCom-zhTW-3.7M.whissle-agent-llm-training-data
Whissle Agent LLM Training Data
Training and validation data for the Whissle Agent LoRA model.
Each sample is a (perception, response) pair where:
Perception = structured ASR output (transcript + emotion + intent + entities + MI behavior)
Response = ideal agent response with SSML prosody, tool calls, MI codes, and reasoning
Dataset Statistics
Split
Samples
Training
5,171
Validation
272
Total
5,443
By Domain
Domain
File… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/whissle-agent-llm-training-data.
