datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IntentClassification_Dataset_for_AI_Assistant_Prompt_Routing_HindiAssistantEval-trajectories
AssistantEval trajectories
This dataset contains recorded AssistantEval runs and their task context. It is published for inspection of completed evaluations. Run records include conversations, service requests and responses, grading evidence, and recorded state. They do not provide executable tasks or resettable service environments.
The source is the public AssistantEval results index, run-data guide, and methodology. The release snapshot was generated October 6, 2026: 1,495… See the full description on the dataset page: https://huggingface.co/datasets/oro-ai/AssistantEval-trajectories.ai-assistant-contributionscomputer_agent_reinforcement_learning_trajectory_seagent_ai_assistant_tools_agent_mcpprompt_engineering_techniques_ai_assistantvoice-assistant-story-qa
Vietnamese Voice Assistant Story QA — Full Audio
60 Vietnamese multiple-choice examples, four stories, one synthetic voice. Three configurations, all using the test split.
Configuration
Skill
Rows
retrieval
Direct information retrieval
20
reasoning
Inference and logic
20
knowledge
Language and general knowledge
20
Input format
Column
Content
audio
Story, question, A–D options and response instruction
transcription
Source text… See the full description on the dataset page: https://huggingface.co/datasets/luvox-ai/voice-assistant-story-qa.Friendly-Pashto-AI-Assistant-Reasoning-Dataset
Friendly Pashto AI Assistant Reasoning Dataset 🚀
دا ډیټاسیټ د iPashto.ai پروژې یوه برخه ده چې د پښتو ژبې د مصنوعي ځیرکتیا (AI) او مرستیالانو د روزنې لپاره په ځانګړي ډول د استدلال (Reasoning) او دوستانه خبرو اترو په موخه جوړ شوی دی.
د کارولو څرنګوالی (How to Load)
تاسو کولی شئ دا ډیټاسیټ په اسانۍ سره د Hugging Face datasets کتابتون له لارې پورته کړئ:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Friendly-Pashto-AI-Assistant-Reasoning-Dataset.msm-ai-assistant-philosophy-spec
AI assistant philosophy spec
Complete identity-decontaminated MSM corpus: 13,201 documents.
Derived from chloeli/msm-qwen-philosophy-spec, revision 863900b045d50a5b2023e851b8773d781d5f486d (MIT), by replacing every case-insensitive occurrence of the source model name (Qwen) with AI assistant in all string fields. All documents, domains, order, and other content are retained. Only text is intended as training input. Provider references and other identity claims have not been… See the full description on the dataset page: https://huggingface.co/datasets/P0u4a/msm-ai-assistant-philosophy-spec.students-coding-questions-from-ai-assistant
Dataset Documentation
Overview
This dataset contains 6776 questions asked by students from CodeAid, an AI coding assistant, during a C programming class over a 12-week semester from January to April 2023. The course did not allow the use of ChatGPT, but CodeAid was permitted. CodeAid, powered by GPT-3, did not directly disclose code solutions even when requested by students. Instead, it functioned like a teaching assistant, providing scaffolded responses in natural… See the full description on the dataset page: https://huggingface.co/datasets/majeedkazemi/students-coding-questions-from-ai-assistant.ai-assistant-hvac-picks
AI assistant picks for HVAC, 2,383 cities
When a homeowner asks an AI assistant "My furnace won't turn on. Who can fix it this week in Plano?", which company does it name, which does it pass over, and why?
This dataset records one AI assistant's answers to six heating and cooling questions in each of 2,383 cities (2,068 US, 88 Canada, 227 UK), September 29 to October 1, 2026: 34,907 picks with the reason given and the page it read, 22,056 companies it considered and passed over… See the full description on the dataset page: https://huggingface.co/datasets/GregM/ai-assistant-hvac-picks.voice-assistant-long-story-qa-stqa
Vietnamese Voice Assistant Long Story QA — STQA
60 multiple-choice examples about two long stories: the injured worker on a train and Oshin’s child at a night-time party. The complete story is shared as text; audio contains each question, its options and response instruction.
Configuration
Skill
Test rows
retrieval
Direct information retrieval (Level 1)
20
reasoning
Inference and logic (Level 2)
20
knowledge
Language and general knowledge (Level 3)
20… See the full description on the dataset page: https://huggingface.co/datasets/luvox-ai/voice-assistant-long-story-qa-stqa.s23-ultra-ai-assistant
S23 Ultra AI Assistant
A fully open Android app that gives a Gemini-powered AI agent direct control of your Samsung Galaxy phone: it sees the screen (screenshot + UI tree), taps, swipes and types for you, launches any installed app, browses the internet, searches the web, downloads files, downloads and installs APKs, sends SMS, places calls, changes volume and opens system settings panels.
Built for the Samsung Galaxy S23 Ultra (works on any Android 13+ phone; tested targets are… See the full description on the dataset page: https://huggingface.co/datasets/dweeeee/s23-ultra-ai-assistant.insurance-ai-assistant-finetuning-datasetsecurity_assistant_dataAssistantEval-results
AssistantEval results
AssistantEval results are aggregate measurements of assistants completing synthetic, stateful tasks. This repository contains the published leaderboard CSV: one row per assistant, with task outcomes, behavior measures, grading counts, confidence intervals, and timing fields.
The CSV is the original public file from assistanteval.com/data/results.csv. The accompanying source index and methodology describe the published snapshot and grading process.… See the full description on the dataset page: https://huggingface.co/datasets/oro-ai/AssistantEval-results.voice-assistant-story-qa-stqa
Vietnamese Voice Assistant Story QA — Story Text / Question Audio
60 Vietnamese multiple-choice examples, four stories, one synthetic voice. Three configurations, all using the test split.
Configuration
Skill
Rows
retrieval
Direct information retrieval
20
reasoning
Inference and logic
20
knowledge
Language and general knowledge
20
Input format
Column
Content
audio
Unchanged question, A–D options and response instruction… See the full description on the dataset page: https://huggingface.co/datasets/luvox-ai/voice-assistant-story-qa-stqa.HuggingChat-AI-Assistants-Deleted-System-Promptsvoice-assistant-story-qa-saqt
Vietnamese Voice Assistant Story QA — Story Audio / Question Text
60 Vietnamese multiple-choice examples, four stories, one synthetic voice. Three configurations, all using the test split.
Configuration
Skill
Rows
retrieval
Direct information retrieval
20
reasoning
Inference and logic
20
knowledge
Language and general knowledge
20
Input format
Column
Content
audio
Story, beginning with Cho câu chuyện sau:
transcription
Source text… See the full description on the dataset page: https://huggingface.co/datasets/luvox-ai/voice-assistant-story-qa-saqt.smolified-eli10-ai-assistant
🤏 smolified-eli10-ai-assistant
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model Snehadev/smolified-eli10-ai-assistant.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 91ccdda8)
Records: 10000
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by Snehadev.
Generated via Smolify.ai.
assistant-ai-agent
Assistant Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP
This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc.
The dataset is helpful for AI… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/assistant-ai-agent.scrum-coaching-assistant
VDF AI — Scrum coaching conversations
English-centric, multi-turn chat examples for supervised fine-tuning (SFT) and alignment work, shared by VDF AI. Conversations frame the assistant as a Scrum and team performance coach drawing on Scrum patterns (in the spirit of ScrumPLoP), systems thinking, and complex adaptive systems perspectives.
Canonical home on Hugging Face: vdf-ai-support/scrum-coaching-assistant.
Clone the dataset repository with Git:
git clone… See the full description on the dataset page: https://huggingface.co/datasets/vdf-ai/scrum-coaching-assistant.AI-doctor-assistant-datasetai-coding-assistants-benchmark-2026
AI Coding Assistants Benchmark 2026 — Methodology Dataset
Independent benchmark methodology for evaluating AI coding assistants in 2026. Covers Claude Code (Anthropic), Cursor, GitHub Copilot, Windsurf (Codeium), Aider, Continue.dev, Cody (Sourcegraph), Tabnine, OpenAI Codex CLI, and Replit Agent.
Methodology
Test bench: 12 real-world coding tasks across Python, TypeScript, Rust, Go
Benchmark: SWE-bench Verified scores per tool (cross-language)
Performance:… See the full description on the dataset page: https://huggingface.co/datasets/Ricco020/ai-coding-assistants-benchmark-2026.business-assistant-ai-toolsAquilaX-AI-security-assistant-reasoning
AquilaX Security Assistant with Reasoning Template
A cybersecurity instruction-tuning dataset converted from AquilaX-AI/security_assistant_data with explicit reasoning template for training models with chain-of-thought capabilities in vulnerability analysis.
Dataset Description
This dataset contains 18,282 examples focused on cybersecurity vulnerability analysis, secure coding practices, and security remediation. Each assistant response includes structured reasoning steps… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/AquilaX-AI-security-assistant-reasoning.ai-engineer-assistant-datasetctrlpotato-ai-interview-assistant-benchmark
CTRLpotato AI Interview Assistant Cross-review Evidence Matrix (2026)
A citation-ready snapshot of hands-on desktop evidence for six AI interview assistants: Cluely, Interview Coder, LockedIn AI, ULTRACODE AI, Parakeet AI, and Final Round AI.
The package contains 66 assessments across 6 products and 11 shared criteria. Product versions and test dates are preserved in every row.
Important scope
This is a cross-review evidence matrix, not a statistically controlled… See the full description on the dataset page: https://huggingface.co/datasets/ae0j/ctrlpotato-ai-interview-assistant-benchmark.SoVTP_pipeline_completed_qwen3-vl-8b-instruct_stage1234only_image_gpt-5-nano_wtSoVTP_pipeline_completed_gpt-5-mini
