Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01transformers-community /circleci-test-resultstextn<1K4 likes5k downloads4mo agoHugging Face02BAAI /CI-VID 📄 CI-VID: A Coherent Interleaved Text-Video Dataset CI-VID is a large-scale dataset designed to advance coherent multi-clip video generation. Unlike traditional text-to-video (T2V) datasets with isolated clip-caption pairs, CI-VID supports text-and-video-to-video (TV2V) generation by providing over 340,000 interleaved sequences of video clips and rich captions. It enables models to learn both intra-clip content and inter-clip transitions, fostering story-driven generation with… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/CI-VID.text100K<n<1M7 likes3.2k downloads10mo agoHugging Face03survivi /grad_cilp0.28_1001M<n<10M0 likes3.2k downloads1y agoHugging Face04evaluate /conll2003-citextn<1K0 likes1.4k downloads4y agoHugging Face05wheres-my-python /floorplans-cityscapes Dataset Summary This is a curated collection of floorplan images sourced from across the internet. It is intended for research in architectural AI, layout generation, and urban scene understanding. Data format: Image files with associated integer labels. Sources: Publicly available images from various web sources (This dataset is one unified collections). Purpose: Educational and research use. Dataset Structure The dataset follows the standard Hugging Face Image… See the full description on the dataset page: https://huggingface.co/datasets/wheres-my-python/floorplans-cityscapes.imagefeature-extraction1K<n<10K1 likes983 downloads7mo agoHugging Face06evaluate /squad-citextn<1K0 likes924 downloads4y agoHugging Face07renzzyyy1028 /civil-code-phil Civilex — Philippine Legal RAG & SFT Dataset Retrieval corpus and supervised fine-tuning (SFT) data for a retrieval-augmented generation (RAG) pipeline over Philippine law: the Civil Code (Republic Act No. 386) and Supreme Court jurisprudence. Produced by the civilex-thesis research pipeline. Contents: 11k+ Supreme Court jurisprudence cases spanning 1949–2025, and 2,270 articles from the Civil Code (Republic Act No. 386). Dataset structure . ├── README.md ├──… See the full description on the dataset page: https://huggingface.co/datasets/renzzyyy1028/civil-code-phil.textquestion-answering100K<n<1M0 likes623 downloads13d agoHugging Face08olm /cia-world-factbook-snapshotstext1K<n<10K1 likes605 downloads4y agoHugging Face09auslawbench /AusLaw-Citation-BenchmarkThis is the dataset proposed in the paper: Methods for Legal Citation Prediction in the Age of LLMs: An Australian Law Case Study. text10K<n<100K3 likes478 downloads1y agoHugging Face10EngineeringAI-LAB /CineBoard3D-plus 🎬 CineBoard3D++: Dynamic 3D Story World Dataset 📊 Dataset Summary CineBoard3D++ is a collection of editable, movie-inspired 3D story worlds built with StoryBlender for narrative-grounded camera planning and world visual attention. It brings together story scripts, animated characters, scene geometry, and shot-level configurations in native Blender projects. The benchmark covers 50 stories, 457 scenes, 1,585 shots, and 3,197 3D assets (836 plot-related and 2,361… See the full description on the dataset page: https://huggingface.co/datasets/EngineeringAI-LAB/CineBoard3D-plus.3dn<1K1 likes462 downloads25d agoHugging Face11cia-tools /parsed_datatext1K<n<10K0 likes456 downloads1y agoHugging Face12KerryMe /cinepile_10ktabular10K<n<100K0 likes439 downloads4mo agoHugging Face13cindyxl /ObjaversePlusPlus Objaverse++: Curated 3D Object Dataset with Quality Annotations Paper Code Chendi Lin, Heshan Liu, Qunshu Lin, Zachary Bright, Shitao Tang, Yihui He, Minghao Liu, Ling Zhu, Cindy Le We cleaned the Objaverse dataset so you don't have to. In this work, we meticulously curated a collection of Objaverse objects and developed an effective classifier capable of scoring the entire Objaverse. Our extensive annotation system considers geometric structure and texture… See the full description on the dataset page: https://huggingface.co/datasets/cindyxl/ObjaversePlusPlus.texttext-to-3d100K<n<1M23 likes417 downloads10mo agoHugging Face14samsepiol4 /netryx-new-york-city-13km Nyc-Core-Usethis 13km Pre-computed MegaLoc index for Netryx Drishti geolocation. Coverage Center: 40.713200, -74.002500 Radius: 13.0 km Panoramas: 663,084 Index entries: 2,652,336 Descriptor model: MegaLoc Descriptor dim: 1024 (PCA from 8448) Usage from netryx_hub import NetryxHub hub = NetryxHub() hub.download("nyc-core-usethis-13km", output_dir="./netryx_data/index") # Now open Netryx and search! Or download manually and use Import Index in… See the full description on the dataset page: https://huggingface.co/datasets/samsepiol4/netryx-new-york-city-13km.tabularn<1K0 likes370 downloads3mo agoHugging Face15opendatalab /CiteVQA CiteVQA English | 简体中文 CiteVQA is a document visual question answering benchmark for faithful evidence attribution. Unlike conventional DocVQA datasets that only score the final answer, CiteVQA requires a model to answer a question with evidence grounded in the source document at the element level. The benchmark is designed to evaluate whether a system can not only answer correctly, but also cite the right supporting region in long, real-world PDFs. The dataset contains 1,897… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/CiteVQA.textvisual-question-answering1K<n<10K11 likes333 downloads5mo agoHugging Face16NetherlandsForensicInstitute /s2orc-citation-pairs-translated-nlThis is a Dutch version of the S2ORC: The Semantic Scholar Open Research Corpus. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. textsentence-similarity10M<n<100M0 likes328 downloads2y agoHugging Face17BigBro23 /CityCube-Benchimagequestion-answering1K<n<10K3 likes315 downloads9mo agoHugging Face18CinderD /TeachArena TeachArena TeachArena is a benchmark for evaluating AI tutoring agents across the full teaching decision chain — from moment-to-moment tutoring dialogue, to pedagogical judgment on packaged evidence, to multi-step teaching workflows grounded in a learning-management system. It contains 354 tasks organized into three stages, a mock LMS environment database, the agent policy documents, and the full scoring logic. Why three stages A capable teaching agent must both… See the full description on the dataset page: https://huggingface.co/datasets/CinderD/TeachArena.texttext-generationn<1K1 likes309 downloads9d agoHugging Face19manus4oHER /cia-readingroom-bucket-00text10K<n<100K0 likes304 downloads3mo agoHugging Face20CinderD /wildtrace WildTrace strict481 WildTrace is a source-internal long-context multi-hop reasoning benchmark built from natural evidence trails. Unlike reverse-synthetic QA, its tasks are mined in situ from long source documents before questions are written. The strict481 release contains 481 locked tasks over 214 public long-form sources, with full-document, evidence-withheld evaluation. The model under test receives only the source document and the public question; evidence spans, clue… See the full description on the dataset page: https://huggingface.co/datasets/CinderD/wildtrace.textquestion-answeringn<1K1 likes298 downloads2mo agoHugging Face21nvidia /Nemotron-RL-Instruction-Following-Citation-Formatting-v1 Dataset Description: Teaches the model to cite specific document parts using reference markers like [ref:1], ref:3, etc. Supports single-reference, multi-reference, and inline citations. This dataset is ready for commercial/non-commercial uses. Dataset Owner(s): NVIDIA Corporation Dataset Creation Date: Created on: April 10, 2026 Last Modified on: April 10, 2026 Version: Nemotron-RL-Instruction-Following-CitationFormatting-v1… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Citation-Formatting-v1.texttext-generation1K<n<10K3 likes284 downloads7d agoHugging Face22csoai /cinematic-world-stills Council of AI — cinematic stills Cinematic stills produced for Council of AI surfaces. metadata.jsonl gives the Hub image viewer a caption per file. These are illustrations — they carry no measurement and back no slot. The live board is the authority GET https://councilof.ai/api/gspc — quote totals.public_count. This Hub card is a printer of that GET, never a second engine. If the fetch fails the honest answer is UNCHECKABLE — never a fabricated 0.000. Status… See the full description on the dataset page: https://huggingface.co/datasets/csoai/cinematic-world-stills.imageothern<1K0 likes250 downloads2h agoHugging Face23commoncrawl /citations Common Crawl Citations Overview This dataset contains citations referencing Common Crawl Foundation and its datasets, pulled from Google Scholar. Please note that these citations are not curated, so they will include some false positives. An annotated subset of these citations with additional fields can be found at cc-citations. text1K<n<10K5 likes249 downloads6mo agoHugging Face24ServiceNow /Dr-CiK Dr-CiK: A Testbed for Foresight-Driven Agents Dr-CiK is a benchmark for evaluating whether agents can retrieve forecasting-relevant context from a noisy document corpus, filter out distractors, distill the retrieved context into forecast-useful evidence, and produce forecasts grounded in that evidence. Real-world time-series forecasting often depends not only on historical observations but also on external context that must be actively discovered from heterogeneous, noisy… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/Dr-CiK.tabulartime-series-forecasting10K<n<100K3 likes244 downloads3mo agoHugging Face25Elfsong /scholar-citation-historytextn<1K1 likes224 downloads2h agoHugging Face26jamescalam /world-cities-geoDataset containing city, country, region, and continents alongside their longitude and latitude co-ordinates. Cartesian coordinates are provided in x, y, z features. tabular1K<n<10K15 likes221 downloads4y agoHugging Face27C-Tianyu /circoimagen<1K0 likes218 downloads6mo agoHugging Face28Jonaszky123 /L-CiteEval L-CITEEVAL: DO LONG-CONTEXT MODELS TRULY LEVERAGE CONTEXT FOR RESPONDING? Paper   Github   Zhihu Benchmark Quickview L-CiteEval is a multi-task long-context understanding with citation benchmark, covering 5 task categories, including single-document question answering, multi-document question answering, summarization, dialogue understanding, and synthetic tasks, encompassing 11 different long-context tasks. The context lengths for these tasks range from 8K to 48K.… See the full description on the dataset page: https://huggingface.co/datasets/Jonaszky123/L-CiteEval.tabularquestion-answering1K<n<10K3 likes209 downloads2y agoHugging Face29opencsg /CIMD CIMD [[中文]] | [[English]] CSGHub Dataset Page | Hugging Face | OpenCSG Community 中文说明 数据集概述 CIMD 是一个面向文档智能任务的跨来源、多语言 JSONL 语料库。当前公开快照包含 111,308 条解析记录,覆盖制度参考、学术与长文档资料、机构分析、企业运营、公共讨论和市场相关材料等来源家族。每条记录都把正文与来源类型、语言、时间、关键词、授权标签和来源字段放在同一个结构里,用户拿到数据后可以直接做检索、抽样、审计和数据治理。 公开数据已转换为统一字段,并按来源家族拆分为可单独加载的子集;它不是原始文件夹的简单打包。用户可以只读取制度参考、学术长文档或公共讨论记录,也可以合并多个子集构建检索库、抽取训练候选样本、构造评测样本池,并按来源、语言和时间字段继续筛选。 CIMD 和通用网页语料的差别在于记录级元数据。它不只提供可索引文本,还提供… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/CIMD.text10K<n<100K7 likes204 downloads4mo agoHugging Face30karrykkk /CI-VID-SC-4072-repro CI-VID SC 4072 Reproduction Subset Reproduction copy of the exact CI-VID subset used by the UniVideo interleaved text/video run: https://wandb.ai/diffusionrl-dyh/univideo-civid/runs/w2292csr Source and license Derived from BAAI/CI-VID, revision 3d6a7b049fd63505a80b751b80d5a0929602ddcf. CI-VID is provided for non-commercial research use. Access to this subset does not replace or broaden the upstream license. Users must comply with the upstream dataset terms.… See the full description on the dataset page: https://huggingface.co/datasets/karrykkk/CI-VID-SC-4072-repro.text1K<n<10K0 likes204 downloads15d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.