Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /laion-tts-annotated-v1 LAION TTS Annotated v1 107,563,551 annotated speech utterances across six subsets — with the audio, the codec tokens and the annotations, all joined by one key. 283,681 audio-hours. Per utterance: the transcript with word-level timings, 40 emotion intensities, 57 VoiceNet voice-character dimensions, four audio-quality heads, vocal-burst detections with timings, and a natural-language caption describing the voice and the delivery — plus the audio itself, its… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-tts-annotated-v1.tabulartext-to-speech100M<n<1B0 likes9.9k downloads15d agoHugging Face02sam-paech /wildchat_creative_writing_annotated_10ktabular10K<n<100K2 likes7.2k downloads9mo agoHugging Face03laion /laion-voice-profiles-annotated Synthetic Voice-Profile Performances Authors: Christoph Schuhmann and LAION. 28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from 500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting conditions. Every utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads, vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a 768-d… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-annotated.tabulartext-to-speech10M<n<100M0 likes4.3k downloads15d agoHugging Face04obadx /muaalem-annotated-v3 قاعدة بيانات المعلم القرآنية هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript وصف قاعدة بيانات العلم مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية البيانات الوصفية للمصاحف ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train'] وصف… See the full description on the dataset page: https://huggingface.co/datasets/obadx/muaalem-annotated-v3.audio100K<n<1M7 likes3.8k downloads1y agoHugging Face05barryallen16 /fitcheck-annotate-datasettext10K<n<100K0 likes3.2k downloads24d agoHugging Face06locuslab /fineweb_annotatedtext100M<n<1B2 likes2.1k downloads5mo agoHugging Face07mlfoundations-dev /r1_annotated_aimetext1K<n<10K0 likes1.9k downloads2y agoHugging Face08mesolitica /Malaysian-Emilia-annotated Malaysian Emilia Annotated Annotate Malaysian-Emilia using Data-Speech pipeline. Malaysian Youtube Originally from malaysia-ai/crawl-youtube Total 3168.8 hours. Gender prediction, filtered-24k_processed_24k_gender.zip Language prediction, filtered-24k_processed_language.zip Force alignment. Post cleaned to 24k and 44k sampling rates, 24k, filtered-24k_processed_24k.zip 44k, filtered-24k_processed_44k.zip Synthetic description… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-annotated.tabulartext-to-speech1M<n<10M2 likes1.7k downloads2y agoHugging Face09smolagents /GAIA-annotateddocumentn<1K1 likes1.6k downloads1y agoHugging Face10laion /openthoughts-4-math-qwen3-32b-7k-annotated-sharegpttext1M<n<10M0 likes1.6k downloads10mo agoHugging Face11marin-community /open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted-selfconsistency Dataset Card for Open-Thoughts-4-30K-Math-Qwen3-32B-Annotated-32768-Tokens-N8-Reformatted-SelfConsistency Overview This dataset is a self-consistency filtered version of marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted. For each prompt, 8 responses were generated by Qwen3-32B with different random seeds. A majority vote was taken over the final answers (extracted from \boxed{...}) to determine the most popular answer, and only… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted-selfconsistency.tabular100K<n<1M4 likes1.4k downloads8mo agoHugging Face12WenqingCao /finevisionmax-annotatedimage10M<n<100M1 likes1.3k downloads4mo agoHugging Face13obadx /mualem-recitations-annotatedaudio100K<n<1M4 likes1.2k downloads1y agoHugging Face14cognaize /elements_annotated_tables_4500_docs Dataset 🚀 Progress Last update (UTC): 2025-11-11 15:40:21Z Documents processed: 4500 / 500058 Batches completed: 30 Total pages/rows uploaded: 89882 Latest batch summary Batch index: 30 Docs in batch: 150 Pages/rows added: 1487 imageobject-detection10K<n<100K0 likes1.2k downloads11mo agoHugging Face15humair025 /Urdu-ONYX-WAV-kanade-Annotated Urdu-ONYX-WAV-real-Annotated Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,217 Total Duration: 42.77 hours Average Duration: 5.87 seconds Duration Range: 0.65s - 122.23s Average Phonemes: 18.5 per sample Average Kanade Tokens: 151.1 per sample Global Embedding Dimension: 128 New Columns This dataset adds the following columns: duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.tabulartext-to-speech100K<n<1M0 likes1.1k downloads8mo agoHugging Face16nour-world /muaalem-annotated-v3 قاعدة بيانات المعلم القرآنية هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript وصف قاعدة بيانات العلم مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية البيانات الوصفية للمصاحف ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train']… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/muaalem-annotated-v3.audio100K<n<1M0 likes965 downloads25d agoHugging Face17HealthDataHub /PARHAF-response_to_treatment-annotated Dataset Card for PARHAF-response_to_treatment-annotated Reporting Issues & Contributing If you encounter any errors or inconsistencies in this dataset, please report them in the discussion section of the "Community" tab on Hugging Face. For more substantial contributions or collaboration opportunities, feel free to contact us directly. Dataset Summary PARHAF-response_to_treatment-annotated is a subpart of the PARHAF corpus, an open French corpus of… See the full description on the dataset page: https://huggingface.co/datasets/HealthDataHub/PARHAF-response_to_treatment-annotated.tabulartoken-classificationn<1K0 likes869 downloads6mo agoHugging Face18OpenMed /synthvision-annotated-qwen synthvision-annotated-qwen Medical images annotated by Qwen 3.5 (397B) via Doubleword Records: 59,476 About First-half annotations from the SynthVision pipeline. 59,476 medical images annotated by Qwen 3.5 (397B MoE, 17B active) via Doubleword batch inference. Each record contains a multi-turn clinical conversation (5-9 turns), a clinical narrative report, structured findings, reasoning chain, and difficulty rating. Schema id: str #… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/synthvision-annotated-qwen.textvisual-question-answering10K<n<100K2 likes814 downloads7mo agoHugging Face19MinhDuk /final_dataset_annotatedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/MinhDuk/final_dataset_annotated.tabularrobotics100K<n<1M2 likes778 downloads8d agoHugging Face20harsh-7070 /COCO-Wholebody-annotatedimage100K<n<1M0 likes662 downloads2y agoHugging Face21marin-community /open-thoughts-4-math-qwen3-32b-annotated Dataset Card for Open-Thoughts-4-Math-Qwen3-32B-Annotated This dataset is the Qwen3-32B annotated version of mlfoundations-dev/hero_run_4_math curated by the OpenThoughts4 team. We provide the responses from Qwen3-32B in the generated_text column. These samples were generated using temperature = 0.8 and max output tokens = 7,500. We note that many of the responses are truncated, so use this dataset wisely! Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-math-qwen3-32b-annotated.tabular1M<n<10M0 likes599 downloads11mo agoHugging Face22Roszczyk /InfraRed_Photos_Annotated_for_Birds_DetectionDataset created using Intel Geti. This dataset was created while working on Master thesis at Warsaw University of Technology. Code and more additional info about the project is available here: DecisionSystemFeeder. The model based on YOLO trained on this dataset using Intel Geti was published here: Infrared_Bird_Detection image0 likes597 downloads26d agoHugging Face23OpenMed /synthvision-annotated-kimi synthvision-annotated-kimi Medical images annotated by Kimi K2.5 via Doubleword Records: 59,539 About Second-half annotations from the SynthVision pipeline. 59,539 medical images annotated by Kimi K2.5 (1T MoE, 32B active) via Doubleword batch inference. Each record contains a multi-turn clinical conversation (5-9 turns), a clinical narrative report, structured findings, reasoning chain, and difficulty rating. Schema id: str # unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/synthvision-annotated-kimi.textvisual-question-answering10K<n<100K1 likes568 downloads7mo agoHugging Face24mlfoundations-dev /openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554 mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces HLE HMMT AIME25 LiveCodeBenchv5 Accuracy 34.3 74.5 79.4 49.4 51.0 44.3 53.9 21.5 23.1 12.2 17.0 22.7 40.1 AIME24 Average Accuracy: 34.33% ± 1.89% Number of Runs: 10 Run Accuracy Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554.tabular10K<n<100K0 likes521 downloads1y agoHugging Face25mlfoundations-dev /math_stratos_scale_judged_and_annotated_with_difficultytabular100K<n<1M0 likes512 downloads2y agoHugging Face26pepijn223 /robocasa_pretrain_human300_v4_annotated5This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "observation.images.robot0_agentview_left": { "dtype": "video", "shape": [ 256, 256, 3 ], "names": [ "height", "width", "channel" ], "video_info": {… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/robocasa_pretrain_human300_v4_annotated5.tabularrobotics10M<n<100M2 likes460 downloads3mo agoHugging Face27soerenray /speech_commands_enriched_and_annotated Dataset Summary 📊 Data-centric AI principles have become increasingly important for real-world use cases.At Renumics we believe that classical benchmark datasets and competitions should be extended to reflect this development. 🔍 This is why we are publishing benchmark datasets with application-specific enrichments (e.g. embeddings, baseline results, uncertainties, label error scores). We hope this helps the ML community in the following ways: Enable new researchers to quickly… See the full description on the dataset page: https://huggingface.co/datasets/soerenray/speech_commands_enriched_and_annotated.audio10K<n<100K2 likes452 downloads3y agoHugging Face28laion /openthoughts-4-code-qwen3-32b-32k-annotatedtabular100K<n<1M2 likes423 downloads10mo agoHugging Face29star092304 /CEFR-Annotated-WordNet CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning Paper: CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning Authors: Masato Kikuchi, Masatsugu Ono, Toshioki Soga, Tetsu Tanabe, Tadachika Ozono Overview CEFR-Annotated WordNet is a comprehensive semantic database that maps WordNet 3.0 senses and SemCor 3.0 instances to CEFR (Common European Framework of Reference for Languages)… See the full description on the dataset page: https://huggingface.co/datasets/star092304/CEFR-Annotated-WordNet.text100K<n<1M1 likes384 downloads4mo agoHugging Face30maximellerbach /omx_multicubes_annotatedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/maximellerbach/omx_multicubes_annotated.tabularrobotics100K<n<1M0 likes355 downloads12d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.