datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
danish-dynaword
🧨 Danish Dynaword
Version
1.2.25 (Changelog)
Language
dan, dansk, Danish
License
Openly Licensed, See the respective dataset
Models
For model trained used this data see danish-foundation-models
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 7.40M
Number of tokens (Llama 3): 9.83B
Average document length in tokens (min, max): 1.33K (2, 19.46M)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-dynaword.swedish-dynaword
🧨 Swedish Dynaword
Version
0.0.13 (Changelog)
Language
Swedish (sv, swe)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 547.06M
Number of tokens (Llama 3): 36.34B
Average document length in tokens (min, max): 66.42 (2, 8.14M)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/swedish-dynaword.multilingual-gsm-symbolic
Multilingual GSM-Symbolic
Version
v0.5.5
Released
2026-10-06
Languages
108 (19 validated)
Multilingual GSM-Symbolic is a benchmark for evaluating arithmetic reasoning in large language models across many languages. It extends Apple's GSM-Symbolic approach by providing symbolic templates that generate thousands of structurally equivalent but numerically distinct math problems. Templates and generation are handled by the multilingual-gsm-symbolic… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/multilingual-gsm-symbolic.norwegian-dynaword
🧨 Norwegian Dynaword
Version
0.0.18 (Changelog)
Language
Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 4.47M
Number of tokens (Llama 3): 9.98B
Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.BKAINewsCorpus
Dataset Card for "BKAINewsCorpus"
The Binhvq News Corpus, a widely used dataset featuring approximately 20 million articles from diverse sources, received its last update in May 2021. To enhance this collection, we gathered an additional 10 million articles up until November 2023. By integrating these newly acquired articles with the existing Binhvq News Corpus, we have created an extensive Vietnamese News Corpus comprising about 32M articles. Subsequent fuzzy deduplication was… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/BKAINewsCorpus.faroese-dyna-instruct
🧨 Faroese dyna-instruct
Version
0.1.1 (Changelog)
Language
Faroese (fao)
License
Openly Licensed, see individual datasets
Models
For models trained on this data see danish-foundation-models
Contact
If you have questions about this project please create an issue here
Dataset Description
Number of samples: 9.18K
Number of tokens (Llama 3): 2.68M
Average conversation length in tokens (min, max): 291.83 (27, 1.24K)
Average number of… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dyna-instruct.faroese-dynaword
🧨 Faroese Dynaword
Version
0.0.8 (Changelog)
Language
Faroese (fo, fao)
License
Openly Licensed, See the respective dataset
Models
Currently there are no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 410.89K
Number of tokens (Llama 3): 59.82M
Average document length in tokens (min, max): 145.6 (2, 208.41K)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dynaword.dutch-dynaword
🧨 Dutch Dynaword
Version
1.0.1 (Changelog)
Language
nld, Nederlands, Dutch
License
Openly Licensed, See the respective dataset
Models
For model trained used this data see danish-foundation-models
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 14.45M
Number of tokens (Llama 3): 37.89B
Average document length in tokens (min, max): 2.62K (2, 5.45M)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dutch-dynaword.icelandic-dynaword
🧨 Icelandic Dynaword
Version
0.0.15 (Changelog)
Language
Icelandic (is, isl)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 39.85M
Number of tokens (Llama 3): 2.67B
Average document length in tokens (min, max): 66.98 (3, 1.03M)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword.multi-ifeval
MultiIFEval
This dataset is an instruction-following dataset for 300+ languages, translated and localised from the English IFEval dataset.
Dataset Details
Dataset Description
All samples come from the English IFEval dataset, and we translate and localise with Gemini-3-flash-preview.
When translating and localising samples, we also include a random Wikipedia article in the target language, both to give some context for localisation, but also to… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/multi-ifeval.danish-gigaword
Danish Gigaword Corpus
Version: 1.0.0
License: See the respective dataset
Dataset Summary
The Danish Gigaword Corpus contains text spanning several domains and forms. This version does not include the sections containing tweets ("General Discussions" and "Parliament Elections"), "danavis", "Common Crawl" and "OpenSubtitles" due to potential privacy, quality and copyright concerns.
Loading the dataset
from datasets import load_dataset
name =… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-gigaword.foundation-models-perturbationData for the paper "Foundation Models Improve Perturbation Response Prediction" as described on GitHub.
NewsSapoVietnamese NewsSapo Dataset
The Vietnamese NewsSapo dataset was constructed to train sentence/passage embeddings. Our dataset is structured in a "title-abstract-contents" format, where each news article is represented by a tuple of (title, abstract, content). The content is the main text body of the article and has been processed to remove images, videos, and other non-textual elements. The dataset contains 31,728,183 triples.
To build this dataset, we followed a two-step process:
Step 1:… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/NewsSapo.DetailCaps-4870
DetailCaps-4870 Benchmark
The detail image caption evaluation benchmark proposed in our paper Benchmarking and Improving Detail Image Caption.
🏠 Homepage | 📑 Paper | 🤗 Huggingface Datasets
Overview
We curate 4870 images from various datasets, accompanying with ground truth detail captions generated by GPT-4V, Gemini-1.5-Pro and GPT-4O for evaluation.
We also provide captions generated by three open-source LVLMs, which are LLaVA-1.5, CogVLM and ShareCaptioner, as well… See the full description on the dataset page: https://huggingface.co/datasets/foundation-multimodal-models/DetailCaps-4870.ai-arenaen
AI-Arenaen: Danish-language conversations and human preferences
AI-Arenaen is a public chatbot arena for Danish users. People chat with two anonymous models side by side and say which answer they prefer. This dataset is the result.
Each row is one turn of a conversation: the two models' answers to the same user message, the preference the user gave on that turn (if any), and the full conversation both answers belong to. The prompts come from real users, most of them writing in… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ai-arenaen.norwegian-dyna-instruct
🧨 Norwegian dyna-instruct
Version
0.1.0 (changelog)
Languages
Norwegian Bokmål (nob), Norwegian Nynorsk (nno), and English (eng) translation input
License
Mixed open licenses; see the table below
Sources
Five datasets (source cards)
Dataset Description
Number of samples: 14.40K
Number of tokens (Llama 3): 6.27M
Average conversation length in tokens (min, max): 435.63 (4, 8.92K)
Average number of turns (min, max): 2.13 (2, 3)… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dyna-instruct.vi-alpaca
🇻🇳 Vietnamese Alpaca Dataset
This dataset is especially designed for Vietnamese based on the idea from Stanford Alpaca and Self-Instruct paper. The motivation behind the creation of this dataset stems from the hope to contribute high-quality dataset to Vietnamese commnunity to train language models.
To construct this dataset, we follow a two-step process:
Step 1: Manually create Vietnamese seed tasks
We employ the methodology outlined in the Self-Instruct paper we meticulously… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/vi-alpaca.dala_gen_v3milp-instances-parquet
MILP instances (Parquet)
Competition-style instances packed as Zstd-compressed Parquet shards for partial downloads.
Schema
Column
Type
Description
instance_id
string
Stem name (e.g. load_balancing_0)
task
string
item_placement, load_balancing, or anonymous
split
string
train or valid
json_text
string
Raw contents of the sidecar .json
mps_gz
binary
Bytes of the .mps.gz file
Tasks are independent (separate folders / configs). Shards are named… See the full description on the dataset page: https://huggingface.co/datasets/foundation-models/milp-instances-parquet.icelandic-dyna-instruct
🧨 Icelandic dyna-instruct
Version
0.1.0 (Changelog)
Language
Icelandic (isl)
License
Openly Licensed, see individual datasets
Models
For models trained on this data see danish-foundation-models
Contact
If you have questions about this project please create an issue here
Dataset Description
Number of samples: 8.11K
Number of tokens (Llama 3): 7.09M
Average conversation length in tokens (min, max): 874.89 (182, 1.39K)
Average number… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dyna-instruct.imagescrosslingual
VNLAWQC, VNSynLawQC: A Vietnamese Legal Retrieval Dataset
VNLAWQC, is sourced from the Vietnamese Law Library (VLL). The VLL contains articles that address questions spanning multiple aspects of the legal domain. Each article provides an answer supported by one or more legal documents, with hyperlinks directing to the corresponding documents.
VNSynLawQC is augmented based on law documents in VNLAWQC using Llama-3-70B.
Dataset Composition
The dataset consists of query… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/crosslingual.dala
DaLA: Danish Linguistic Acceptability Evaluation Dataset
DaLA (paper) is a benchmark dataset for linguistic acceptability judgment in Danish, designed to evaluate how well NLP models, especially large language models (LLMs), understand grammaticality in real-world Danish sentences. The dataset extends previous resources by introducing a broader and more realistic set of error types and providing data splits suitable for evaluation via few-shot or finetuning.
🔗… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dala.ifeval-da
IFEval-da
This dataset is a translation of the English IFEval dataset,
which was published in this paper and contains 541 prompts,
each with a combination of one or more of 25 different constraints. The dataset was professionally
translated and localised by expert native speakers.
Dataset Details
Translated by: Rasmus Larsen (rasmus.larsen@alexandra.dk), Nathalie Hau Sørensen (naha@hum.ku.dk) and Kenneth Enevoldsen (kenneth.enevoldsen@cas.au.dk)
Funded by: Danish… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ifeval-da.ai-arenaen-raw
AI-Arenaen: Danish-language conversations and human preferences
AI-Arenaen is a public chatbot arena for Danish users. People chat with two anonymous models side by side and say which answer they prefer. This dataset is the result.
Each row is one turn of a conversation: the two models' answers to the same user message, the preference the user gave on that turn (if any), and the full conversation both answers belong to. The prompts come from real users, most of them writing in… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ai-arenaen-raw.golden-batch-sentinel-data
Golden Batch Sentinel Data
Benchmark datasets for process monitoring and fault detection in batch manufacturing.
Datasets
IndPenSim (Industrial Penicillin Simulation)
A 100,000L fermentation simulation with 100 batches and rich multivariate signals.
Source: Mendeley Data
Paper: Modern day monitoring and control challenges...
Batches: 100 (90 normal, 10 faulty)
Variables: 37 process variables (Raman spectra excluded for efficiency)
Time resolution: 0.2 hours… See the full description on the dataset page: https://huggingface.co/datasets/foundation-models/golden-batch-sentinel-data.pi-rsi-experiment-artifactsvietnamese-roleplay-realm
🇻🇳 Vietnamese Role-play Realm Dataset
This is a dataset of GPT-generated Vietnamese characters made to increase the ability of open-source language models to role-play.
It contains 446 characters generated by GPT-3.5
Each character will have 20 topics generated by ChatGPT. And each topic will have a conversation corresponding with it
In 446 characters, there are 400 general characters and 46 Vietnamese characters.
To construct this dataset, we follow a four-step process:… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/vietnamese-roleplay-realm.synthetic-values-model-charter
value_units.jsonl is the individual parsed values from the model charter.
scenarios.jsonl is invididual hypothetical scenarios based on the values in values_units.jsonl
sft_*.jsonl generated accepted responses.
dpo_*.jsonl generated accepted+rejected responses.
vi-self-chat-sharegpt-format
🇻🇳 Vietnamese Self-Chat Dataset
This dataset is designed to enhance the model's ability to engage in multi-turn conversations with humans.
To construct this dataset, we follow a two-step process:
Step 1: Instruction Generation
We employ the methodology outlined in the Self-Instruct paper to craft a diverse set of instructions. This paper serves as a guide for aligning pretrained language models with specific instructions, providing a structured foundation for subsequent dialogue… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/vi-self-chat-sharegpt-format.
