Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yatin-superintelligence /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M53 likes4.2k downloads7mo agoHugging Face02RUC-NLPIR /Omnimodal-Agent-SFT-2K OmniGAIA: Omni-Modal General AI Assistant Benchmark 📄 Paper   •   💻 Code & Demo   •   🤗 Dataset & Model   •   📈 Leaderboard This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.audioquestion-answering1K<n<10K9 likes2.2k downloads7mo agoHugging Face03yatin-superintelligence /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M29 likes1.3k downloads7mo agoHugging Face04ziggylott /agent-tts-libraryaudion<1K0 likes793 downloads2mo agoHugging Face05leungtianle /AgentChataudio100K<n<1M2 likes754 downloads6mo agoHugging Face06Jaward /lectura-agents-data LectūraAgents Dataset Overview This dataset is in support of findings in our paper LectūraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching. LectūaAgents is a hierarchical multi-agent framework that enables end-to-end personalized learning experiences through adaptive embodied teaching. It mirrors a professor–students’ relationship, wherein a ProfessorAgent guides a collaborative team of specialized subordinate… See the full description on the dataset page: https://huggingface.co/datasets/Jaward/lectura-agents-data.audion<1K24 likes610 downloads1mo agoHugging Face07BlueIsGreen /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/BlueIsGreen/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M11 likes610 downloads7mo agoHugging Face08yatin-superintelligence /Audio-Video-Engineering-Agentic-Tasks-1M Audio/Video Engineering Agentic Tasks (1M) Abstract A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.tabulartext-generation1M<n<10M14 likes599 downloads7mo agoHugging Face09voidful /agent-sft-stitch-zh-tts agent-sft-stitch-zh-tts Voiced version of voidful/agent-sft-stitch-zh: the STITCH-S spoken chunks synthesized with BlueMagpie-TTS (hung_yi_lee voice), per-utterance loudness-aligned to -23 LUFS, best-of-N + Whisper-CER accepted. Configs records (default): one row per agent dialogue — id/source/user/msg (full STITCH-S trajectory) + available_tools + STITCH quality scores + spoken (ordered list of the utterances, each with audio, text, seg_index, cer, accepted… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts.audiotext-to-speech100K<n<1M0 likes592 downloads3mo agoHugging Face10rAVEUK /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M3 likes554 downloads7mo agoHugging Face11leungtianle /AgentChat-Test Test Set Description This directory contains the test set used for tool-use evaluation. The JSON files under Test-JSON/ are organized by task type: SingleTaskProcessing/tool-select_test.json: single-tool selection tasks. ParallelProcessing/parallel-call_test.json: parallel tool-call tasks. ProactiveSeeking/searchTools_test_predictions_kept.json: proactive tool-search tasks. TaskDecomposition/muti-tool-select_test.json: multi-tool task decomposition tasks.… See the full description on the dataset page: https://huggingface.co/datasets/leungtianle/AgentChat-Test.audion<1K0 likes504 downloads3mo agoHugging Face12kryp1234 /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M1 likes474 downloads7mo agoHugging Face13Torenn /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/Torenn/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M1 likes470 downloads7mo agoHugging Face14cx-cmu /AgentWebBench-corpus AgentWebBench Corpus Pre-built dense-retrieval corpus for AgentWebBench [ICML 2026], a benchmark for Multi-Agent Coordination in Agentic Web over a realistic 100-website slice of ClueWeb22 (~18.4M documents). This repository holds the embeddings and FAISS indices the benchmark loads at run time, including per-website indices, a global index, and website-level vectors. It does not contain ClueWeb22 text (see Raw documents). Websites: 100 Documents: ~18.4M Embedding dim: 1024… See the full description on the dataset page: https://huggingface.co/datasets/cx-cmu/AgentWebBench-corpus.audiotext-retrievaln<1K0 likes396 downloads4mo agoHugging Face15ppenner /edge-agent-reasoning-websearch-260k Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/ppenner/edge-agent-reasoning-websearch-260k.texttext-generation100K<n<1M1 likes367 downloads4mo agoHugging Face16DEMIRUNC /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes272 downloads7mo agoHugging Face17JACKYS999 /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/JACKYS999/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes222 downloads5mo agoHugging Face18kanepi-1977 /Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/kanepi-1977/Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes200 downloads7mo agoHugging Face19ArkhAngelLifeJiggy /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes160 downloads12d agoHugging Face20svryn /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/svryn/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes133 downloads5mo agoHugging Face21tterumiimurett1 /agentic-asrgated Agentic ASR Public consolidated audio and ASR result dataset for the OSWorld and WildClawBench benchmark families. Layout osworld/: synthetic raw/colloquial speech, human recordings, DNS-noise pairs, task images, ASR results, and reports. wildclawbench/: 60 formal colloquialized prompts, synthetic speech, 20 synthetic ASR condition tables, and ten-participant human recordings. task0_template derivatives are excluded. metadata/conditions.jsonl: model, variant… See the full description on the dataset page: https://huggingface.co/datasets/tterumiimurett1/agentic-asr.audio10K<n<100K0 likes131 downloads19d agoHugging Face22jatshi /trusted-full-duplex-agent-data Trusted Full-Duplex Speech Agent — Training Data & Evidence Companion dataset for model: jatshi/trusted-full-duplex-agent and GitHub: Jatshi/trusted-full-duplex-agent. 2.0 evidence update The 2.0 release adds the evidence needed to reproduce the deployed runtime rather than only the offline research pipeline: real-time latency telemetry, unified guardrail voice samples, the exact turn-taking MLP reports, TrustGate/ASR runtime configuration, and integrity… See the full description on the dataset page: https://huggingface.co/datasets/jatshi/trusted-full-duplex-agent-data.audioaudio-text-to-textn<1K0 likes115 downloads21d agoHugging Face23person9601 /agentCourse_storagestorage for gaia question files audion<1K1 likes113 downloads1y agoHugging Face24AustinZhang /resp-agent-dataset Resp-229K: Respiratory Sound Dataset A Large-Scale Respiratory Sound Dataset for Training and Evaluation 📖 Overview Resp-229K is a comprehensive respiratory sound dataset containing 229,101 valid audio files with a total duration of over 407 hours. This dataset is curated for training the Resp-Agent system - an intelligent respiratory sound analysis and generation framework. 📊 Dataset Statistics Split Valid Files Total Duration Avg Duration Max… See the full description on the dataset page: https://huggingface.co/datasets/AustinZhang/resp-agent-dataset.audioaudio-classification100K<n<1M0 likes105 downloads8mo agoHugging Face25voidful /agent-sft-stitch-zh-tts-sampleaudio1K<n<10K0 likes30 downloads3mo agoHugging Face26ronith09 /Agent-TTS-cleanedaudio10K<n<100K0 likes28 downloads1y agoHugging Face27AgentPublic /eval-stt-officiels EvalSTT — Corpus officiels (FR) Corpus public d'évaluation de modèles de transcription (speech-to-text) sur du langage de l'administration française : discours officiels, allocutions publiques et questions au gouvernement. Constitué par le département IA dans l'État (DINUM) dans le cadre de l'évaluation des modèles de speech-to-text. Ce dataset est publié pour la transparence : il documente les jeux de données utilisés pour nos évaluations et permet de reproduire les mesures… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/eval-stt-officiels.audioautomatic-speech-recognitionn<1K1 likes27 downloads4mo agoHugging Face28leungtianle /agent-chataudio1K<n<10K0 likes24 downloads11mo agoHugging Face29AgenticFilmLab /kandinsky-6-native-video-audio-test-record Episode 009 — Kandinsky 6.0: native video and audio test record Watch the published episode: Kandinsky 6.0 Makes Its Own Sound. Does It Follow the Prompt? 31 complete outputs shown in Agentic Film Lab's episode, including the official samurai and dragon examples, recurring characters, physical actions, six visual styles, controlled prompting follow-ups, and four quantized tests. This is a dataset of experiment evidence, not model weights or a trained model. What is… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFilmLab/kandinsky-6-native-video-audio-test-record.videon<1K0 likes24 downloads2d agoHugging Face30voidful /agent-sft-stitch-zh-tts-taste-codec-sample agent-sft-stitch-zh-tts Taste-S codec sample Ten accepted synthesized clips sampled from voidful/agent-sft-stitch-zh-tts, encoded with andybi7676/taste-s-en-zhtw-small-gemma4. Extraction follows IntelliGen's stage1_extract_taste.py: the 24 kHz source audio is resampled to 16 kHz and converted to 80-bin cool-whisper features. The encoder is conditioned on the external transcript tokenized with the Gemma 4 tokenizer, with streaming disabled. codec_indices has shape… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-sample.audioaudio-to-audion<1K0 likes23 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.