Team Ai
3 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01phucpx247 /turn-detection-vietnameseNguồn dữ liệu: vi-wiki-conversational-search Tỷ lệ Complete:Incomplete = 244304:366451 Đã lưu 610755 samples vào training_data.csv Đã lưu 6189 samples vào test_data.csv texttext-classification100K<n<1M4 likes100 downloads1y agoHugging Face02justpluso /turn_detection_3k_zh这是一个专为儿童机器人聊天场景设计的语料数据集,由 Gemini 2.5 Pro 生成。数据集旨在充分模拟儿童的聊天行为,覆盖了广泛的主题和对话类型,包括单轮和三轮对话。 主要涵盖的场景类别包括: 学科与知识探索:语文、数学、英语、科学、历史、地理及常识问答。 文化与传统:民俗节日、神话传说、礼仪习惯。 兴趣爱好与玩乐:玩具、各类游戏(电子、棋类、户外)、收藏。 创意表达与想象:绘画手工、音乐歌舞、故事创作、幻想世界。 个人与社交生活:关于自己、家庭亲人、朋友同学、学校生活(非学术方面)。 日常生活与环境观察:天气季节、饮食食物、动植物观察、交通工具、周围环境事件。 媒体与娱乐:动画片、漫画绘本、电影、儿童歌曲故事音频、适龄网络内容。 对机器人的互动与探索:询问机器人基本信息、能力功能、情感互动、测试挑战及音量调节等。 涵盖以下场景 A. 学科与知识探索 (Academic & Knowledge Exploration) 语文 (Chinese Language Arts): 认字、写字、组词、造句 古诗词(背诵、含义、诗人故事) 成语(含义、故事、接龙)… See the full description on the dataset page: https://huggingface.co/datasets/justpluso/turn_detection_3k_zh.texttext-classification1K<n<10K5 likes59 downloads1y agoHugging Face03PuristanLabs1 /Urdu-Turn-Detection-10k Urdu Turn Detection Dataset 🗣️ A high-quality dataset of 10,000 Urdu sentences labeled for Turn Detection (End-of-Turn). This dataset is designed to help conversational AI systems determine if a user has finished speaking (Complete) or is pausing/trailing off (Incomplete). Dataset Details Total Samples: 10,000 Language: Urdu (ur) - Nastaliq/Arabic Script only. Cleanliness: - 100% Urdu Script (No Roman/English). Avg. Sentence Length: - ~7.7 words (33 characters)… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/Urdu-Turn-Detection-10k.texttext-classification10K<n<100K0 likes20 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.