datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
goldsrc-models-datasetmodels-logodanish-dynaword
🧨 Danish Dynaword
Version
1.2.25 (Changelog)
Language
dan, dansk, Danish
License
Openly Licensed, See the respective dataset
Models
For model trained used this data see danish-foundation-models
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 7.40M
Number of tokens (Llama 3): 9.83B
Average document length in tokens (min, max): 1.33K (2, 19.46M)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-dynaword.open-models-prompt-datasets
🖼️ Open Models Prompt Dataset
🖼️ The ultimate open models image prompt dataset (10GB+). 5400+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for various open-source AI image models and the resulting generated images. The entire dataset exceeds 10GB and contains 5400+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/open-models-prompt-datasets.swedish-dynaword
🧨 Swedish Dynaword
Version
0.0.13 (Changelog)
Language
Swedish (sv, swe)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 547.06M
Number of tokens (Llama 3): 36.34B
Average document length in tokens (min, max): 66.42 (2, 8.14M)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/swedish-dynaword.Hoyoverse_Character_Modelsrvc-modelsCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
norwegian-dynaword
🧨 Norwegian Dynaword
Version
0.0.18 (Changelog)
Language
Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 4.47M
Number of tokens (Llama 3): 9.98B
Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.faroese-dynaword
🧨 Faroese Dynaword
Version
0.0.8 (Changelog)
Language
Faroese (fo, fao)
License
Openly Licensed, See the respective dataset
Models
Currently there are no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 410.89K
Number of tokens (Llama 3): 59.82M
Average document length in tokens (min, max): 145.6 (2, 208.41K)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dynaword.dutch-dynaword
🧨 Dutch Dynaword
Version
1.0.1 (Changelog)
Language
nld, Nederlands, Dutch
License
Openly Licensed, See the respective dataset
Models
For model trained used this data see danish-foundation-models
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 14.45M
Number of tokens (Llama 3): 37.89B
Average document length in tokens (min, max): 2.62K (2, 5.45M)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dutch-dynaword.vrm-premium-modelsicelandic-dynaword
🧨 Icelandic Dynaword
Version
0.0.15 (Changelog)
Language
Icelandic (is, isl)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 39.85M
Number of tokens (Llama 3): 2.67B
Average document length in tokens (min, max): 66.98 (3, 1.03M)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword.cgaxis-3d-models-sample
CGAxis 3D Models - Free Sample (Furniture / Chairs)
A free, licensed sample of human-authored 3D models from CGAxis, a 3D content studio operating since 2008. This sample is a taster of the full CGAxis AI Data corpus (4,200+ 3D models + 7,913 PBR material sets) available for commercial AI-training licenses.
Every model ships as GLB and USDZ (the USDZ with UsdPhysics authored: rigid body, collision, mass, physics material), with geometry statistics, real-world scale in… See the full description on the dataset page: https://huggingface.co/datasets/CGAxis/cgaxis-3d-models-sample.modelsModelSafetyBenchmodelscar-modelsIllustrius.modelsnorwegian-dyna-instruct
🧨 Norwegian dyna-instruct
Version
0.1.0 (changelog)
Languages
Norwegian Bokmål (nob), Norwegian Nynorsk (nno), and English (eng) translation input
License
Mixed open licenses; see the table below
Sources
Five datasets (source cards)
Dataset Description
Number of samples: 14.40K
Number of tokens (Llama 3): 6.27M
Average conversation length in tokens (min, max): 435.63 (4, 8.92K)
Average number of turns (min, max): 2.13 (2, 3)… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dyna-instruct.HEV-ORF1-models
Hepatitis E virus ORF1 (nsp1) — AlphaFold2 model collection
1,178 AlphaFold2 predictions of the HEV ORF1 (nsp1) replicase, packaged so a
static web app can render the 3D model, the predicted aligned error (PAE) matrix and the multiple
sequence alignment without a server.
open the viewer: https://tubiana.github.io/ORF1viewer (this dataset is its data root)
repository — app + pipeline code, no data: https://github.com/tubiana/tubiana.github.io
dataset repo:… See the full description on the dataset page: https://huggingface.co/datasets/ttubiana/HEV-ORF1-models.diffusion_models_course_stickerspmf-concurrency-models
Concurrency-aware process model forecasting: models and drawings
Weekly process models for the BPI2017, BPI2019 and Hospital Billing event logs, produced
by the code of the paper
Concurrency-Aware Process Model Forecasting with Causal Nets,
and a drawing of each one as a workflow net. The code is at
github.com/YongboYu/pmf-concurrency.
Contents
models/<tag>/<Log>_pmf/weekly_models/window_<n>_<family>.json causal nets
models/<tag>/<Log>_pmf/structural_metrics.csv… See the full description on the dataset page: https://huggingface.co/datasets/pmf-liris/pmf-concurrency-models.Spatial-Blind-Spots-in-Vision-Language-Modelslicense: mit
model_evaluated:
name: Qwen3-VL-2B-Instruct
url: https://huggingface.co/Qwen/Qwen3-VL-2B-Instruct
evaluation_notebook:
https://www.kaggle.com/code/wajidhassanmoosa/blind-spot-qwen3-2b
evaluation_setup: |
The model evaluated in this study is Qwen3-VL-2B-Instruct.
Evaluation was conducted using the Hugging Face Transformers library
with automatic device mapping (device_map="auto") and "bfloat16" dtype selection.
For each example:
The image was provided as part of a… See the full description on the dataset page: https://huggingface.co/datasets/hassan-wajid/Spatial-Blind-Spots-in-Vision-Language-Models.stable_diffusion_modelsimagesA_Reasoning_Critique_of_Diffusion_Models
A Reasoning Critique of Diffusion Models
Author: Zixi "Oz" Li (李籽溪)
Date: December 12, 2025
Type: Theoretical AI Research (Geometry, Reasoning Theory)
Citation
@misc{oz_lee_2025,
author = { Oz Lee },
title = { A_Reasoning_Critique_of_Diffusion_Models (Revision 267326d) },
year = 2025,
url = { https://huggingface.co/datasets/OzTianlu/A_Reasoning_Critique_of_Diffusion_Models },
doi = { 10.57967/hf/7243 }… See the full description on the dataset page: https://huggingface.co/datasets/OzTianlu/A_Reasoning_Critique_of_Diffusion_Models.Long-he-mineru-models
English | 简体中文
🚀Access MinerU Now→✅ Zero-Install Web Version ✅ Full-Featured Desktop Client ✅ Instant API Access; Skip deployment headaches – get all product formats in one click. Developers, dive in!
👋 join us on Discord and WeChat
MinerU — High-accuracy document parsing engine for LLM · RAG · Agent workflows
Converts PDF · DOCX · PPTX · XLSX · Images · Web pages into structured Markdown / JSON · VLM+OCR dual engine · 109 languages
MCP Server ·… See the full description on the dataset page: https://huggingface.co/datasets/Lh23593217/Long-he-mineru-models.vastu-modelsmatryoshka-diffusion-models-paper-examples
Matryoshka Diffusion Models - paper examples
This dataset contains the 1024x1024 images included in the Matryoshka Diffusion Models
paper.
Arxiv: https://arxiv.org/abs/2310.15111
SD.models
