Team Ai
15 results

Data annotation

MCG-NJU /VideoChat3-Training-Data-Annotations VideoChat3-Stage3-Training-Data This repository includes all annotation files used across the four training stages of VideoChat3, from Stage 0 to Stage 3. You can refer to the provided source-data links to download videos, images, and other multimedia data for training. In videochat3_data_annotations, we also provide a source field to indicate the source dataset for each entry. To facilitate Stage 3 training reproduction using the high-quality open-source datasets we collected… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-Training-Data-Annotations.video-text-to-text2 likes654 downloads8d agoHugging Facechfeng /probe-tip-annotations-data0 likes509 downloads6mo agoHugging Facemarijanic /map-annotation-tool-data Map Annotation Tool Media This dataset contains Docling-extracted figure crops grouped by source PDF. It is the media input for map-annotation-tool; human annotations are maintained separately in the application repository and submitted through pull requests. Contents media/figures/brgm-v1/: 1,512 figures from the earlier BRGM parsing batch, of which 156 had legacy annotations at migration time. media/figures/brgm-v2/: 6,046 figures from the newer BRGM parsing… See the full description on the dataset page: https://huggingface.co/datasets/marijanic/map-annotation-tool-data.imageimage-classification1K<n<10K1 likes293 downloads1mo agoHugging Facejasongraf1 /annotation_app_data Dataset Card for Systematic Review of Acceptability Judgments data A curated dataset of research articles used in a systematic review of judgment tasks in linguistics. Each entry records article-level metadata and experiment-level methodological features, supporting structured comparison and analysis across studies. Dataset Description This annotation dataset comprises systematically coded observations from a corpus of published studies employing judgment tasks in… See the full description on the dataset page: https://huggingface.co/datasets/jasongraf1/annotation_app_data.tabularn<1K0 likes279 downloads3d agoHugging Facegatilin /LocateAnything-Data-ShareGPT-Annotation LocateAnything Full Hosted ShareGPT 41 hosted views, 88,508,086 ShareGPT records, approximately 28 GiB of JSONL annotations. This package represents the entire hosted training pool, not the earlier capped 30M tokenized cache. Contents And Rights data/: messages with role / content and portable locanytar://pool/member image references. dataset_info.json: LlamaFactory registration including ShareGPT tags. initial_adapter/: previous LocateAnything 27k-step LoRA used… See the full description on the dataset page: https://huggingface.co/datasets/gatilin/LocateAnything-Data-ShareGPT-Annotation.object-detection10M<n<100M0 likes228 downloads11h agoHugging FaceTTS-AGI /voice-annotation-data-v2 Voice Annotation Data v2 A curated dataset of 18,632 audio samples (9,391 positives + 9,241 negatives) across 58 voice dimensions. Each bucket contains up to 25 positive examples (audio that clearly fits the bucket) and 25 negative examples (audio confirmed to NOT fit the bucket by Gemini 2.0 Flash). Changes from v1 Positive + Negative pairs: Every bucket now has up to 25 confirmed negative examples alongside 25 positives EXPL redefined: Content Appropriateness reduced… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/voice-annotation-data-v2.audioaudio-classification10K<n<100K3 likes164 downloads6mo agoHugging Face