Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01banned-historical-archives /banned-historical-archives 和谐历史档案馆数据集 - Banned Historical Archives Datasets 和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。 目录结构 banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步 raw # 原始文件 config # 配置文件 todo # 存放暂未录入网站的文件 部分报纸和图片资料存放在单独的仓库: 名称 地址 状态 参考消息 https://huggingface.co/datasets/banned-historical-archives/ckxx 未录入 人民日报 https://huggingface.co/datasets/banned-historical-archives/rmrb 已精选重要的文章录入 文汇报… See the full description on the dataset page: https://huggingface.co/datasets/banned-historical-archives/banned-historical-archives.imagen<1K94 likes1.2m downloads1y agoHugging Face02picbreeder-vlm /picbreeder-vlm-archive Picbreeder-VLM Archive Every image evolved by the swarm of vision-language-model "breeders" in In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models (GECCO 2026), together with the CPPN genomes that produced them, the agents' reasoning transcripts, the lineage graphs, and the analysis artifacts behind the paper and blog. The original Picbreeder (Secretan et al., 2008) let crowds of humans collaboratively evolve images from CPPN… See the full description on the dataset page: https://huggingface.co/datasets/picbreeder-vlm/picbreeder-vlm-archive.imageimage-to-text100K<n<1M14 likes44k downloads3mo agoHugging Face03Dragonegg2026 /banned-historical-archives 和谐历史档案馆数据集 - Banned Historical Archives Datasets 和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。 目录结构 banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步 raw # 原始文件 config # 配置文件 todo # 存放暂未录入网站的文件 部分报纸和图片资料存放在单独的仓库: 名称 地址 状态 参考消息 https://huggingface.co/datasets/banned-historical-archives/ckxx 未录入 人民日报 https://huggingface.co/datasets/banned-historical-archives/rmrb 已精选重要的文章录入 文汇报… See the full description on the dataset page: https://huggingface.co/datasets/Dragonegg2026/banned-historical-archives.imagen<1K1 likes4k downloads6mo agoHugging Face04Memento-ARC /ultimatedocument0 likes4k downloads2mo agoHugging Face05Augmentiv /ArchiveDatadocumentn<1K0 likes2.8k downloads6mo agoHugging Face06diwash-barla /meta-archiveimage1K<n<10K1 likes2.4k downloads2d agoHugging Face07eastbrush /eastbrush_archive Eastbrush Archive Official Website (Full Archive System): https://www.eastbrush.com This dataset contains high-resolution images and structured tags for AI training. The full archive system — including chapter exhibitions, structural context, and extended records — is available on the official website. What is my true self? The Eastbrush Archive is a long-term, evolving system that documents the visual language ofJang Byeong Eun (Eastbrush / 張炳彥) — a painter whose… See the full description on the dataset page: https://huggingface.co/datasets/eastbrush/eastbrush_archive.image1K<n<10K2 likes2.3k downloads7d agoHugging Face08OmniAICreator /ASMR-Archive-Processed ASMR-Archive-Processed (WIP) Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended. Work in Progress — expect breaking changes while the pipeline and data layout stabilize. This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.imageautomatic-speech-recognition98 likes2.2k downloads6mo agoHugging Face09Vasy7777 /cs2-demo-archive CS2 Demo to Dataset — Pipeline Output Samples Sample archives produced by the open-source cs2-demo-to-dataset pipeline, which converts a single CS2 .dem replay file into per-round, per-player first-person video aligned to tick-level state, input and event tables. This release is not a dataset contribution. The point of the upload is to demonstrate that the pipeline produces a coherent, reproducible archive format. Please see the GitHub repository for the recorder code… See the full description on the dataset page: https://huggingface.co/datasets/Vasy7777/cs2-demo-archive.imageother1K<n<10K2 likes1.5k downloads4mo agoHugging Face10Archatext /AraMS-Restore AraMS-Restore — Real Damaged Arabic Manuscript Lines 177 line images cropped from real damaged pages of a historical Arabic manuscript (book_09), each with its transcription. This is the evaluation input for AraMS-Restore: the restoration models are trained on synthetic degradation, and these lines are the honest test of whether that transfers to genuine manuscript decay. There are no clean counterparts and no ground-truth restored images — the damage is what was on the page.… See the full description on the dataset page: https://huggingface.co/datasets/Archatext/AraMS-Restore.imageimage-to-imagen<1K0 likes1.2k downloads2mo agoHugging Face11NameFrame /arcades-unreal-synthetic Arcade Town: a synthetic detection capture rendered in Unreal Engine 300 rendered frames from a single Unreal Engine environment, carrying 5,374 labelled object instances. Every box and every mask here is read out of the engine's own per-instance ID buffer at render time. No model produced these labels and no one drew them by hand, so a label is wrong only where the scene description behind it is wrong. The capture Engine… See the full description on the dataset page: https://huggingface.co/datasets/NameFrame/arcades-unreal-synthetic.imageobject-detectionn<1K0 likes1.2k downloads12d agoHugging Face12aaaad1 /banned-historical-archives 和谐历史档案馆数据集 - Banned Historical Archives Datasets 和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。 目录结构 banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步 raw # 原始文件 config # 配置文件 todo # 存放暂未录入网站的文件 部分报纸和图片资料存放在单独的仓库: 名称 地址 状态 参考消息 https://huggingface.co/datasets/banned-historical-archives/ckxx 未录入 人民日报 https://huggingface.co/datasets/banned-historical-archives/rmrb 已精选重要的文章录入 文汇报… See the full description on the dataset page: https://huggingface.co/datasets/aaaad1/banned-historical-archives.imagen<1K0 likes1k downloads5mo agoHugging Face13banned-historical-archives /hkgongshangribaoimage1K<n<10K0 likes892 downloads2y agoHugging Face14hmar-heritage-org /corpus-archivegated corpus-archive [!WARNING] Experimental Dataset Architecture: The repository structure, metadata tiers, category taxonomies, and catalog indexing formats are currently under active design and evaluation. All specifications, metadata keys, and JSON schemas detailed below represent representational examples and intended targets. This repository serves as a structured digital textual archive preserving Hmar literature, historical accounts, school textbooks, dictionaries, parallel… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/corpus-archive.imagetext-classificationn<1K4 likes889 downloads25d agoHugging Face15karo2w /archiveimagen<1K0 likes821 downloads1y agoHugging Face16banned-historical-archives /hkgongshangwanbaoimage1K<n<10K0 likes808 downloads2y agoHugging Face17pnsk-lab /scratch-archiveimage3 likes800 downloads7mo agoHugging Face18benlehrburger /modern-architectureimage1K<n<10K4 likes761 downloads3y agoHugging Face19banned-historical-archives /dagongbaoimage1K<n<10K0 likes759 downloads2y agoHugging Face20ebylmz /architectsimage1K<n<10K0 likes690 downloads2y agoHugging Face21ritwika96 /arc-agi-3-full-curriculum ARC-AGI-3 full transformation curriculum This release contains 444,000 deterministic multimodal episodes across five curriculum configs: stateful: 168,000 episodes over 12 stateful causal families; search: 76,000 episodes over 6 graph, search, and objective families; physics: 96,000 episodes over 8 object-physics families. composition: 64,000 episodes over 8 pairwise mechanic templates, with complete validation and test pairings absent from training. ambiguity: 40,000 episodes… See the full description on the dataset page: https://huggingface.co/datasets/ritwika96/arc-agi-3-full-curriculum.imageimage-text-to-text100K<n<1M1 likes639 downloads19d agoHugging Face22Reasat /arch_gastric Dataset Card for "arch_gastric" More Information needed image100K<n<1M0 likes601 downloads4y agoHugging Face23ArchaeonSeq /nanochat nanochat nanochat is the simplest experimental harness for training LLMs. It is designed to run on a single GPU node, the code is minimal/hackable, and it covers all major LLM stages including tokenization, pretraining, finetuning, evaluation, inference, and a chat UI. For example, you can train your own GPT-2 capability LLM (which cost $43,000 to train in 2019) for only $48 (2 hours of 8XH100 GPU node) and then talk to it in a familiar ChatGPT-like web UI. On a spot instance… See the full description on the dataset page: https://huggingface.co/datasets/ArchaeonSeq/nanochat.imagen<1K0 likes592 downloads1mo agoHugging Face24banned-historical-archives /huaqiaoribaoimage1K<n<10K0 likes589 downloads2y agoHugging Face25taechasith /kala-uap-archivedocumentn<1K0 likes579 downloads5mo agoHugging Face26aarhus-city-archives /historical-danish-handwriting Dataset Description Dataset Summary The Historical Danish handwriting dataset is a Danish-language dataset containing more than 11.000 pages of transcribed and proofread handwritten text. The dataset currently consists of the published minutes from a number of City and Parish Council meetings, all dated between 1841 and 1939. Languages All the text is in Danish. The BCP-47 code for Danish is da. Dataset Structure Data Instances Each data… See the full description on the dataset page: https://huggingface.co/datasets/aarhus-city-archives/historical-danish-handwriting.imageimage-to-text1K<n<10K4 likes536 downloads1y agoHugging Face27previtus /jpl_trace_gases_archivegeospatialn<1K0 likes505 downloads24d agoHugging Face28archya /MathNet Quick Start · Overview · Tasks · Comparison · Dataset Stats · Data Sources · Pipeline · Schema · License · Citation This is the official MathNet v0. A larger version v1 will be uploaded soon (more countires, problems and richer metadata). Schema is stable but field values may be revised in v1. Quick start from datasets import load_dataset # Default: all problems ds = load_dataset("ShadenA/MathNet", split="train") # Or a specific country / competition-body config arg… See the full description on the dataset page: https://huggingface.co/datasets/archya/MathNet.imagequestion-answering10K<n<100K0 likes464 downloads5mo agoHugging Face29aditya487 /cbi-archive-raw Central Bank of Ireland Archive: original source files 6,309 original files, 6.56 GB. Every PDF, spreadsheet, Word document and archive gathered from the Central Bank of Ireland's public website, stored by content hash so that a search result can be turned back into the document a human would actually read. This is the raw tier. If you want the text, you almost certainly want aditya487/cbi-archive-corpus instead: 5,568 documents and 89,242 page or pseudo-page rows as Parquet… See the full description on the dataset page: https://huggingface.co/datasets/aditya487/cbi-archive-raw.document1K<n<10K0 likes436 downloads1mo agoHugging Face30Arthur12137 /SoftVTBench-archive SoftVTBench — archive Frozen snapshot of everything that lived in Arthur12137/SoftVTBench before the 2026-08-13 re-release: the evaluation USD assets, the soft-body assets, and the first partial data drops. This repo is not maintained. The current dataset is at Arthur12137/SoftVTBench. Contents: eval-assets/, soft-assets/, object-rigid/, object-soft/, spatial-rigid/, spatial-soft/ (9319 files, 2.3 GB). Original dataset card (kept verbatim) SoftVTBench dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arthur12137/SoftVTBench-archive.imageroboticsn<1K0 likes433 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.