datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lasa1m-annotate-part-12lasa1m-annotate-part-15lasa1m-annotate-part-16lasa1m-annotate-part-07lasa1m-annotate-part-06lasa1m-annotate-part-14lasa1m-annotate-part-09lasa1m-annotate-part-13lasa1m-annotate-part-05lasa1m-annotate-part-10lasa1m-annotate-part-11lasa1m-annotate-part-08pico-robotics-annotated
This edition has been merged into the Advanced Edition
The Annotated Edition is discontinued. Its time-stamped action captions (segments.json) are now
included in every episode of the Advanced Edition.
Free entry point: Basic Edition
Point clouds, MCAP and captions: Advanced Edition
Commercial licensing and the full 10,000+ hour collection: jiuchen@openelephant.ai
— OpenElephant Intelligence (公象智能)
wildchat_creative_writing_annotated_10klaion-tts-annotated-v1
LAION TTS Annotated v1
107,563,551 annotated speech utterances across six subsets — with the audio, the codec tokens
and the annotations, all joined by one key.
283,681 audio-hours. Per utterance: the transcript with word-level timings, 40 emotion
intensities, 57 VoiceNet voice-character dimensions, four audio-quality heads, vocal-burst
detections with timings, and a natural-language caption describing the voice and the
delivery — plus the audio itself, its… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-tts-annotated-v1.laion-voice-profiles-annotated
Synthetic Voice-Profile Performances
Authors: Christoph Schuhmann and LAION.
28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from
500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting
conditions. Every
utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads,
vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a
768-d… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-annotated.muaalem-annotated-v3
قاعدة بيانات المعلم القرآنية
هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء التجويد عن طريق رسم صوتي يصف كل قواعد التجويد وصفات الحروف: quran-trainscript
وصف قاعدة بيانات العلم
مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية
البيانات الوصفية للمصاحف
ds = load_dataset('obadx/muaalem-annotated-v3', name='moshaf_metadata')['train']
وصف… See the full description on the dataset page: https://huggingface.co/datasets/obadx/muaalem-annotated-v3.fitcheck-annotate-datasetlasa1m-annotate-part-04lasa1m-annotate-part-01lasa1m-annotate-part-03r1_annotated_aimelasa1m-annotate-part-02dronescapes2_annotated_train_set
Dataset Card for DroneScapes2 (annotated train set)
This is a FiftyOne dataset with 218 samples. It's a subset of this split from the original repo.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/dronescapes2_annotated_train_set.Malaysian-Emilia-annotated
Malaysian Emilia Annotated
Annotate Malaysian-Emilia using Data-Speech pipeline.
Malaysian Youtube
Originally from malaysia-ai/crawl-youtube
Total 3168.8 hours.
Gender prediction, filtered-24k_processed_24k_gender.zip
Language prediction, filtered-24k_processed_language.zip
Force alignment.
Post cleaned to 24k and 44k sampling rates,
24k, filtered-24k_processed_24k.zip
44k, filtered-24k_processed_44k.zip
Synthetic description… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-annotated.openthoughts-4-math-qwen3-32b-7k-annotated-sharegptGAIA-annotatedopen-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted-selfconsistency
Dataset Card for Open-Thoughts-4-30K-Math-Qwen3-32B-Annotated-32768-Tokens-N8-Reformatted-SelfConsistency
Overview
This dataset is a self-consistency filtered version of marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted. For each prompt, 8 responses were generated by Qwen3-32B with different random seeds. A majority vote was taken over the final answers (extracted from \boxed{...}) to determine the most popular answer, and only… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted-selfconsistency.mualem-recitations-annotatedfinevisionmax-annotated
