Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hampta /goldsrc-models-datasetimage10K<n<100K0 likes54k downloads1mo agoHugging Face02huggingchat /models-logoimagen<1K5 likes19k downloads1y agoHugging Face03danish-foundation-models /danish-dynaword 🧨 Danish Dynaword Version 1.2.25 (Changelog) Language dan, dansk, Danish License Openly Licensed, See the respective dataset Models For model trained used this data see danish-foundation-models Contact If you have question about this project please create an issue here Dataset Description Number of samples: 7.40M Number of tokens (Llama 3): 9.83B Average document length in tokens (min, max): 1.33K (2, 19.46M) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-dynaword.imagetext-generation10M<n<100M23 likes10k downloads9d agoHugging Face04Goku-OpenLab /open-models-prompt-datasets 🖼️ Open Models Prompt Dataset 🖼️ The ultimate open models image prompt dataset (10GB+). 5400+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for various open-source AI image models and the resulting generated images. The entire dataset exceeds 10GB and contains 5400+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/open-models-prompt-datasets.image1K<n<10K2 likes3.5k downloads3mo agoHugging Face05danish-foundation-models /swedish-dynaword 🧨 Swedish Dynaword Version 0.0.13 (Changelog) Language Swedish (sv, swe) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 547.06M Number of tokens (Llama 3): 36.34B Average document length in tokens (min, max): 66.42 (2, 8.14M) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/swedish-dynaword.imagetext-generation1B<n<10B3 likes2.3k downloads1mo agoHugging Face06BlenderResources /Hoyoverse_Character_Modelsimagen<1K0 likes1.8k downloads1mo agoHugging Face07AI-C /rvc-modelsCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference imagen<1K1 likes1.7k downloads3y agoHugging Face08danish-foundation-models /norwegian-dynaword 🧨 Norwegian Dynaword Version 0.0.18 (Changelog) Language Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 4.47M Number of tokens (Llama 3): 9.98B Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.imagetext-generation10M<n<100M7 likes1.4k downloads1mo agoHugging Face09danish-foundation-models /faroese-dynaword 🧨 Faroese Dynaword Version 0.0.8 (Changelog) Language Faroese (fo, fao) License Openly Licensed, See the respective dataset Models Currently there are no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 410.89K Number of tokens (Llama 3): 59.82M Average document length in tokens (min, max): 145.6 (2, 208.41K) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dynaword.imagetext-generation1M<n<10M3 likes1.1k downloads10d agoHugging Face10danish-foundation-models /dutch-dynaword 🧨 Dutch Dynaword Version 1.0.1 (Changelog) Language nld, Nederlands, Dutch License Openly Licensed, See the respective dataset Models For model trained used this data see danish-foundation-models Contact If you have question about this project please create an issue here Dataset Description Number of samples: 14.45M Number of tokens (Llama 3): 37.89B Average document length in tokens (min, max): 2.62K (2, 5.45M) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dutch-dynaword.imagetext-generation10M<n<100M3 likes1k downloads1mo agoHugging Face11ayyappanallamothu4 /vrm-premium-modelsimagen<1K2 likes974 downloads5mo agoHugging Face12danish-foundation-models /icelandic-dynaword 🧨 Icelandic Dynaword Version 0.0.15 (Changelog) Language Icelandic (is, isl) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 39.85M Number of tokens (Llama 3): 2.67B Average document length in tokens (min, max): 66.98 (3, 1.03M) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword.imagetext-generation100M<n<1B4 likes901 downloads1mo agoHugging Face13CGAxis /cgaxis-3d-models-sample CGAxis 3D Models - Free Sample (Furniture / Chairs) A free, licensed sample of human-authored 3D models from CGAxis, a 3D content studio operating since 2008. This sample is a taster of the full CGAxis AI Data corpus (4,200+ 3D models + 7,913 PBR material sets) available for commercial AI-training licenses. Every model ships as GLB and USDZ (the USDZ with UsdPhysics authored: rigid body, collision, mass, physics material), with geometry statistics, real-world scale in… See the full description on the dataset page: https://huggingface.co/datasets/CGAxis/cgaxis-3d-models-sample.3dimage-to-3dn<1K0 likes808 downloads4d agoHugging Face14Ouzhang /modelsimagen<1K0 likes690 downloads4mo agoHugging Face15ssu-csec /ModelSafetyBenchdocument10K<n<100K0 likes421 downloads6mo agoHugging Face16Kizi-Art /modelsimagen<1K0 likes367 downloads3y agoHugging Face17natsu39 /car-modelsimage10K<n<100K1 likes252 downloads1y agoHugging Face18LeeAeron /Illustrius.modelsimagen<1K0 likes237 downloads8mo agoHugging Face19danish-foundation-models /norwegian-dyna-instruct 🧨 Norwegian dyna-instruct Version 0.1.0 (changelog) Languages Norwegian Bokmål (nob), Norwegian Nynorsk (nno), and English (eng) translation input License Mixed open licenses; see the table below Sources Five datasets (source cards) Dataset Description Number of samples: 14.40K Number of tokens (Llama 3): 6.27M Average conversation length in tokens (min, max): 435.63 (4, 8.92K) Average number of turns (min, max): 2.13 (2, 3)… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dyna-instruct.imagequestion-answering10K<n<100K0 likes202 downloads1mo agoHugging Face20ttubiana /HEV-ORF1-models Hepatitis E virus ORF1 (nsp1) — AlphaFold2 model collection 1,178 AlphaFold2 predictions of the HEV ORF1 (nsp1) replicase, packaged so a static web app can render the 3D model, the predicted aligned error (PAE) matrix and the multiple sequence alignment without a server. open the viewer: https://tubiana.github.io/ORF1viewer (this dataset is its data root) repository — app + pipeline code, no data: https://github.com/tubiana/tubiana.github.io dataset repo:… See the full description on the dataset page: https://huggingface.co/datasets/ttubiana/HEV-ORF1-models.imageother1K<n<10K0 likes179 downloads1mo agoHugging Face21mariarivaille /diffusion_models_course_stickersimagen<1K1 likes174 downloads2y agoHugging Face22pmf-liris /pmf-concurrency-models Concurrency-aware process model forecasting: models and drawings Weekly process models for the BPI2017, BPI2019 and Hospital Billing event logs, produced by the code of the paper Concurrency-Aware Process Model Forecasting with Causal Nets, and a drawing of each one as a workflow net. The code is at github.com/YongboYu/pmf-concurrency. Contents models/<tag>/<Log>_pmf/weekly_models/window_<n>_<family>.json causal nets models/<tag>/<Log>_pmf/structural_metrics.csv… See the full description on the dataset page: https://huggingface.co/datasets/pmf-liris/pmf-concurrency-models.image1K<n<10K0 likes172 downloads12d agoHugging Face23hassan-wajid /Spatial-Blind-Spots-in-Vision-Language-Modelslicense: mit model_evaluated: name: Qwen3-VL-2B-Instruct url: https://huggingface.co/Qwen/Qwen3-VL-2B-Instruct evaluation_notebook: https://www.kaggle.com/code/wajidhassanmoosa/blind-spot-qwen3-2b evaluation_setup: | The model evaluated in this study is Qwen3-VL-2B-Instruct. Evaluation was conducted using the Hugging Face Transformers library with automatic device mapping (device_map="auto") and "bfloat16" dtype selection. For each example: The image was provided as part of a… See the full description on the dataset page: https://huggingface.co/datasets/hassan-wajid/Spatial-Blind-Spots-in-Vision-Language-Models.imagen<1K7 likes148 downloads7mo agoHugging Face24az8720255 /stable_diffusion_modelsimagen<1K2 likes146 downloads3y agoHugging Face25foundation-models /imagesimagen<1K0 likes145 downloads2mo agoHugging Face26OzTianlu /A_Reasoning_Critique_of_Diffusion_Models A Reasoning Critique of Diffusion Models Author: Zixi "Oz" Li (李籽溪) Date: December 12, 2025 Type: Theoretical AI Research (Geometry, Reasoning Theory) Citation @misc{oz_lee_2025, author = { Oz Lee }, title = { A_Reasoning_Critique_of_Diffusion_Models (Revision 267326d) }, year = 2025, url = { https://huggingface.co/datasets/OzTianlu/A_Reasoning_Critique_of_Diffusion_Models }, doi = { 10.57967/hf/7243 }… See the full description on the dataset page: https://huggingface.co/datasets/OzTianlu/A_Reasoning_Critique_of_Diffusion_Models.documentn<1K3 likes115 downloads10mo agoHugging Face27Lh23593217 /Long-he-mineru-models English | 简体中文 🚀Access MinerU Now→✅ Zero-Install Web Version ✅ Full-Featured Desktop Client ✅ Instant API Access; Skip deployment headaches – get all product formats in one click. Developers, dive in! 👋 join us on Discord and WeChat MinerU — High-accuracy document parsing engine for LLM · RAG · Agent workflows Converts PDF · DOCX · PPTX · XLSX · Images · Web pages into structured Markdown / JSON · VLM+OCR dual engine · 109 languages MCP Server ·… See the full description on the dataset page: https://huggingface.co/datasets/Lh23593217/Long-he-mineru-models.documentn<1K0 likes112 downloads5mo agoHugging Face28Leabert7x /vastu-models3dn<1K0 likes102 downloads15d agoHugging Face29multimodalart /matryoshka-diffusion-models-paper-examples Matryoshka Diffusion Models - paper examples This dataset contains the 1024x1024 images included in the Matryoshka Diffusion Models paper. Arxiv: https://arxiv.org/abs/2310.15111 imagen<1K0 likes92 downloads3y agoHugging Face30LeeAeron /SD.modelsimagen<1K0 likes90 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.