Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nightgoodl /lasa1m-annotate-part-121 likes208k downloads2mo agoHugging Face02nightgoodl /lasa1m-annotate-part-150 likes88k downloads2mo agoHugging Face03nightgoodl /lasa1m-annotate-part-160 likes88k downloads2mo agoHugging Face04nightgoodl /lasa1m-annotate-part-070 likes86k downloads2mo agoHugging Face05nightgoodl /lasa1m-annotate-part-06image100K<n<1M0 likes83k downloads2mo agoHugging Face06nightgoodl /lasa1m-annotate-part-14image100K<n<1M0 likes82k downloads2mo agoHugging Face07nightgoodl /lasa1m-annotate-part-090 likes79k downloads2mo agoHugging Face08nightgoodl /lasa1m-annotate-part-130 likes74k downloads2mo agoHugging Face09nightgoodl /lasa1m-annotate-part-050 likes67k downloads2mo agoHugging Face10nightgoodl /lasa1m-annotate-part-100 likes57k downloads2mo agoHugging Face11nightgoodl /lasa1m-annotate-part-111 likes55k downloads2mo agoHugging Face12nightgoodl /lasa1m-annotate-part-080 likes40k downloads2mo agoHugging Face13skycn110 /pico-robotics-annotatedgated This edition has been merged into the Advanced Edition The Annotated Edition is discontinued. Its time-stamped action captions (segments.json) are now included in every episode of the Advanced Edition. Free entry point: Basic Edition Point clouds, MCAP and captions: Advanced Edition Commercial licensing and the full 10,000+ hour collection: jiuchen@openelephant.ai — OpenElephant Intelligence (公象智能) 3 likes7.6k downloads1m agoHugging Face14sam-paech /wildchat_creative_writing_annotated_10ktabular10K<n<100K2 likes7.1k downloads9mo agoHugging Face15laion /laion-tts-annotated-v1 LAION TTS Annotated v1 107,563,551 annotated speech utterances across six subsets — with the audio, the codec tokens and the annotations, all joined by one key. 283,681 audio-hours. Per utterance: the transcript with word-level timings, 40 emotion intensities, 57 VoiceNet voice-character dimensions, four audio-quality heads, vocal-burst detections with timings, and a natural-language caption describing the voice and the delivery — plus the audio itself, its… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-tts-annotated-v1.tabulartext-to-speech100M<n<1B0 likes4.9k downloads15d agoHugging Face16laion /laion-voice-profiles-annotated Synthetic Voice-Profile Performances Authors: Christoph Schuhmann and LAION. 28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from 500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting conditions. Every utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads, vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a 768-d… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-annotated.tabulartext-to-speech10M<n<100M0 likes4.6k downloads15d agoHugging Face17obadx /muaalem-annotated-v3 قاعدة بيانات المعلم القرآنية هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript وصف قاعدة بيانات العلم مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية البيانات الوصفية للمصاحف ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train'] وصف… See the full description on the dataset page: https://huggingface.co/datasets/obadx/muaalem-annotated-v3.audio100K<n<1M7 likes3.8k downloads1y agoHugging Face18barryallen16 /fitcheck-annotate-datasettext10K<n<100K0 likes3.4k downloads23d agoHugging Face19nightgoodl /lasa1m-annotate-part-040 likes2.4k downloads2mo agoHugging Face20nightgoodl /lasa1m-annotate-part-010 likes2.2k downloads2mo agoHugging Face21nightgoodl /lasa1m-annotate-part-030 likes2k downloads2mo agoHugging Face22mlfoundations-dev /r1_annotated_aimetext1K<n<10K0 likes2k downloads2y agoHugging Face23nightgoodl /lasa1m-annotate-part-02image100K<n<1M0 likes1.8k downloads2mo agoHugging Face24Voxel51 /dronescapes2_annotated_train_set Dataset Card for DroneScapes2 (annotated train set) This is a FiftyOne dataset with 218 samples. It's a subset of this split from the original repo. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/dronescapes2_annotated_train_set.imageimage-classification1K<n<10K2 likes1.8k downloads11mo agoHugging Face25mesolitica /Malaysian-Emilia-annotated Malaysian Emilia Annotated Annotate Malaysian-Emilia using Data-Speech pipeline. Malaysian Youtube Originally from malaysia-ai/crawl-youtube Total 3168.8 hours. Gender prediction, filtered-24k_processed_24k_gender.zip Language prediction, filtered-24k_processed_language.zip Force alignment. Post cleaned to 24k and 44k sampling rates, 24k, filtered-24k_processed_24k.zip 44k, filtered-24k_processed_44k.zip Synthetic description… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-annotated.tabulartext-to-speech1M<n<10M2 likes1.7k downloads2y agoHugging Face26laion /openthoughts-4-math-qwen3-32b-7k-annotated-sharegpttext1M<n<10M0 likes1.7k downloads10mo agoHugging Face27smolagents /GAIA-annotateddocumentn<1K1 likes1.6k downloads1y agoHugging Face28marin-community /open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted-selfconsistency Dataset Card for Open-Thoughts-4-30K-Math-Qwen3-32B-Annotated-32768-Tokens-N8-Reformatted-SelfConsistency Overview This dataset is a self-consistency filtered version of marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted. For each prompt, 8 responses were generated by Qwen3-32B with different random seeds. A majority vote was taken over the final answers (extracted from \boxed{...}) to determine the most popular answer, and only… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted-selfconsistency.tabular100K<n<1M4 likes1.5k downloads8mo agoHugging Face29obadx /mualem-recitations-annotatedaudio100K<n<1M4 likes1.4k downloads1y agoHugging Face30WenqingCao /finevisionmax-annotatedimage10M<n<100M1 likes1.3k downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.