Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01danish-foundation-models /danish-dynaword 🧨 Danish Dynaword Version 1.2.25 (Changelog) Language dan, dansk, Danish License Openly Licensed, See the respective dataset Models For model trained used this data see danish-foundation-models Contact If you have question about this project please create an issue here Dataset Description Number of samples: 7.40M Number of tokens (Llama 3): 9.83B Average document length in tokens (min, max): 1.33K (2, 19.46M) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-dynaword.imagetext-generation10M<n<100M23 likes10k downloads8d agoHugging Face02danish-foundation-models /swedish-dynaword 🧨 Swedish Dynaword Version 0.0.13 (Changelog) Language Swedish (sv, swe) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 547.06M Number of tokens (Llama 3): 36.34B Average document length in tokens (min, max): 66.42 (2, 8.14M) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/swedish-dynaword.imagetext-generation1B<n<10B3 likes2.3k downloads1mo agoHugging Face03danish-foundation-models /multilingual-gsm-symbolic Multilingual GSM-Symbolic Version v0.5.5 Released 2026-10-06 Languages 108 (19 validated) Multilingual GSM-Symbolic is a benchmark for evaluating arithmetic reasoning in large language models across many languages. It extends Apple's GSM-Symbolic approach by providing symbolic templates that generate thousands of structurally equivalent but numerically distinct math problems. Templates and generation are handled by the multilingual-gsm-symbolic… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/multilingual-gsm-symbolic.text100K<n<1M3 likes1.8k downloads4d agoHugging Face04danish-foundation-models /norwegian-dynaword 🧨 Norwegian Dynaword Version 0.0.18 (Changelog) Language Norwegian (no, nor), including Bokmål (nb, nob) and Nynorsk (nn, nno) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 4.47M Number of tokens (Llama 3): 9.98B Average document length in tokens (min… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dynaword.imagetext-generation10M<n<100M7 likes1.4k downloads1mo agoHugging Face05bkai-foundation-models /BKAINewsCorpus Dataset Card for "BKAINewsCorpus" The Binhvq News Corpus, a widely used dataset featuring approximately 20 million articles from diverse sources, received its last update in May 2021. To enhance this collection, we gathered an additional 10 million articles up until November 2023. By integrating these newly acquired articles with the existing Binhvq News Corpus, we have created an extensive Vietnamese News Corpus comprising about 32M articles. Subsequent fuzzy deduplication was… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/BKAINewsCorpus.text10M<n<100M14 likes1.4k downloads3y agoHugging Face06danish-foundation-models /faroese-dyna-instruct 🧨 Faroese dyna-instruct Version 0.1.1 (Changelog) Language Faroese (fao) License Openly Licensed, see individual datasets Models For models trained on this data see danish-foundation-models Contact If you have questions about this project please create an issue here Dataset Description Number of samples: 9.18K Number of tokens (Llama 3): 2.68M Average conversation length in tokens (min, max): 291.83 (27, 1.24K) Average number of… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dyna-instruct.texttext-generation10K<n<100K2 likes1.2k downloads5d agoHugging Face07danish-foundation-models /faroese-dynaword 🧨 Faroese Dynaword Version 0.0.8 (Changelog) Language Faroese (fo, fao) License Openly Licensed, See the respective dataset Models Currently there are no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 410.89K Number of tokens (Llama 3): 59.82M Average document length in tokens (min, max): 145.6 (2, 208.41K) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dynaword.imagetext-generation1M<n<10M3 likes1.1k downloads10d agoHugging Face08danish-foundation-models /dutch-dynaword 🧨 Dutch Dynaword Version 1.0.1 (Changelog) Language nld, Nederlands, Dutch License Openly Licensed, See the respective dataset Models For model trained used this data see danish-foundation-models Contact If you have question about this project please create an issue here Dataset Description Number of samples: 14.45M Number of tokens (Llama 3): 37.89B Average document length in tokens (min, max): 2.62K (2, 5.45M) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dutch-dynaword.imagetext-generation10M<n<100M3 likes1k downloads1mo agoHugging Face09danish-foundation-models /icelandic-dynaword 🧨 Icelandic Dynaword Version 0.0.15 (Changelog) Language Icelandic (is, isl) License Openly Licensed, See the respective dataset Models Currently there is no models trained on this dataset Contact If you have question about this project please create an issue here Dataset Description Number of samples: 39.85M Number of tokens (Llama 3): 2.67B Average document length in tokens (min, max): 66.98 (3, 1.03M) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword.imagetext-generation100M<n<1B4 likes901 downloads1mo agoHugging Face10danish-foundation-models /multi-ifeval MultiIFEval This dataset is an instruction-following dataset for 300+ languages, translated and localised from the English IFEval dataset. Dataset Details Dataset Description All samples come from the English IFEval dataset, and we translate and localise with Gemini-3-flash-preview. When translating and localising samples, we also include a random Wikipedia article in the target language, both to give some context for localisation, but also to… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/multi-ifeval.text100K<n<1M2 likes791 downloads3mo agoHugging Face11danish-foundation-models /danish-gigaword Danish Gigaword Corpus Version: 1.0.0 License: See the respective dataset Dataset Summary The Danish Gigaword Corpus contains text spanning several domains and forms. This version does not include the sections containing tweets ("General Discussions" and "Parliament Elections"), "danavis", "Common Crawl" and "OpenSubtitles" due to potential privacy, quality and copyright concerns. Loading the dataset from datasets import load_dataset name =… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/danish-gigaword.texttext-generation100K<n<1M9 likes357 downloads2y agoHugging Face12genbio-ai /foundation-models-perturbationData for the paper "Foundation Models Improve Perturbation Response Prediction" as described on GitHub. text100K<n<1M0 likes309 downloads8mo agoHugging Face13bkai-foundation-models /NewsSapoVietnamese NewsSapo Dataset The Vietnamese NewsSapo dataset was constructed to train sentence/passage embeddings. Our dataset is structured in a "title-abstract-contents" format, where each news article is represented by a tuple of (title, abstract, content). The content is the main text body of the article and has been processed to remove images, videos, and other non-textual elements. The dataset contains 31,728,183 triples. To build this dataset, we followed a two-step process: Step 1:… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/NewsSapo.textsummarization1M<n<10M6 likes302 downloads3y agoHugging Face14foundation-multimodal-models /DetailCaps-4870 DetailCaps-4870 Benchmark The detail image caption evaluation benchmark proposed in our paper Benchmarking and Improving Detail Image Caption. 🏠 Homepage | 📑 Paper | 🤗 Huggingface Datasets Overview We curate 4870 images from various datasets, accompanying with ground truth detail captions generated by GPT-4V, Gemini-1.5-Pro and GPT-4O for evaluation. We also provide captions generated by three open-source LVLMs, which are LLaVA-1.5, CogVLM and ShareCaptioner, as well… See the full description on the dataset page: https://huggingface.co/datasets/foundation-multimodal-models/DetailCaps-4870.text1K<n<10K15 likes270 downloads2y agoHugging Face15danish-foundation-models /ai-arenaen AI-Arenaen: Danish-language conversations and human preferences AI-Arenaen is a public chatbot arena for Danish users. People chat with two anonymous models side by side and say which answer they prefer. This dataset is the result. Each row is one turn of a conversation: the two models' answers to the same user message, the preference the user gave on that turn (if any), and the full conversation both answers belong to. The prompts come from real users, most of them writing in… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ai-arenaen.texttext-generation1K<n<10K1 likes259 downloads11h agoHugging Face16danish-foundation-models /norwegian-dyna-instruct 🧨 Norwegian dyna-instruct Version 0.1.0 (changelog) Languages Norwegian Bokmål (nob), Norwegian Nynorsk (nno), and English (eng) translation input License Mixed open licenses; see the table below Sources Five datasets (source cards) Dataset Description Number of samples: 14.40K Number of tokens (Llama 3): 6.27M Average conversation length in tokens (min, max): 435.63 (4, 8.92K) Average number of turns (min, max): 2.13 (2, 3)… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/norwegian-dyna-instruct.imagequestion-answering10K<n<100K0 likes202 downloads1mo agoHugging Face17bkai-foundation-models /vi-alpaca 🇻🇳 Vietnamese Alpaca Dataset This dataset is especially designed for Vietnamese based on the idea from Stanford Alpaca and Self-Instruct paper. The motivation behind the creation of this dataset stems from the hope to contribute high-quality dataset to Vietnamese commnunity to train language models. To construct this dataset, we follow a two-step process: Step 1: Manually create Vietnamese seed tasks We employ the methodology outlined in the Self-Instruct paper we meticulously… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/vi-alpaca.text10K<n<100K25 likes189 downloads3y agoHugging Face18danish-foundation-models /dala_gen_v3text1K<n<10K0 likes178 downloads6mo agoHugging Face19foundation-models /milp-instances-parquet MILP instances (Parquet) Competition-style instances packed as Zstd-compressed Parquet shards for partial downloads. Schema Column Type Description instance_id string Stem name (e.g. load_balancing_0) task string item_placement, load_balancing, or anonymous split string train or valid json_text string Raw contents of the sidecar .json mps_gz binary Bytes of the .mps.gz file Tasks are independent (separate folders / configs). Shards are named… See the full description on the dataset page: https://huggingface.co/datasets/foundation-models/milp-instances-parquet.text10K<n<100K0 likes174 downloads6mo agoHugging Face20danish-foundation-models /icelandic-dyna-instruct 🧨 Icelandic dyna-instruct Version 0.1.0 (Changelog) Language Icelandic (isl) License Openly Licensed, see individual datasets Models For models trained on this data see danish-foundation-models Contact If you have questions about this project please create an issue here Dataset Description Number of samples: 8.11K Number of tokens (Llama 3): 7.09M Average conversation length in tokens (min, max): 874.89 (182, 1.39K) Average number… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dyna-instruct.texttext-generation10K<n<100K1 likes174 downloads1mo agoHugging Face21foundation-models /imagesimagen<1K0 likes145 downloads2mo agoHugging Face22bkai-foundation-models /crosslingual VNLAWQC, VNSynLawQC: A Vietnamese Legal Retrieval Dataset VNLAWQC, is sourced from the Vietnamese Law Library (VLL). The VLL contains articles that address questions spanning multiple aspects of the legal domain. Each article provides an answer supported by one or more legal documents, with hyperlinks directing to the corresponding documents. VNSynLawQC is augmented based on law documents in VNLAWQC using Llama-3-70B. Dataset Composition The dataset consists of query… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/crosslingual.feature-extraction3 likes137 downloads2y agoHugging Face23danish-foundation-models /dala DaLA: Danish Linguistic Acceptability Evaluation Dataset DaLA (paper) is a benchmark dataset for linguistic acceptability judgment in Danish, designed to evaluate how well NLP models, especially large language models (LLMs), understand grammaticality in real-world Danish sentences. The dataset extends previous resources by introducing a broader and more realistic set of error types and providing data splits suitable for evaluation via few-shot or finetuning. 🔗… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/dala.texttext-classification1K<n<10K0 likes107 downloads8mo agoHugging Face24danish-foundation-models /ifeval-da IFEval-da This dataset is a translation of the English IFEval dataset, which was published in this paper and contains 541 prompts, each with a combination of one or more of 25 different constraints. The dataset was professionally translated and localised by expert native speakers. Dataset Details Translated by: Rasmus Larsen (rasmus.larsen@alexandra.dk), Nathalie Hau Sørensen (naha@hum.ku.dk) and Kenneth Enevoldsen (kenneth.enevoldsen@cas.au.dk) Funded by: Danish… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ifeval-da.texttext-generationn<1K1 likes93 downloads8mo agoHugging Face25danish-foundation-models /ai-arenaen-rawgated AI-Arenaen: Danish-language conversations and human preferences AI-Arenaen is a public chatbot arena for Danish users. People chat with two anonymous models side by side and say which answer they prefer. This dataset is the result. Each row is one turn of a conversation: the two models' answers to the same user message, the preference the user gave on that turn (if any), and the full conversation both answers belong to. The prompts come from real users, most of them writing in… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ai-arenaen-raw.texttext-generation10K<n<100K0 likes85 downloads11h agoHugging Face26foundation-models /golden-batch-sentinel-data Golden Batch Sentinel Data Benchmark datasets for process monitoring and fault detection in batch manufacturing. Datasets IndPenSim (Industrial Penicillin Simulation) A 100,000L fermentation simulation with 100 batches and rich multivariate signals. Source: Mendeley Data Paper: Modern day monitoring and control challenges... Batches: 100 (90 normal, 10 faulty) Variables: 37 process variables (Raman spectra excluded for efficiency) Time resolution: 0.2 hours… See the full description on the dataset page: https://huggingface.co/datasets/foundation-models/golden-batch-sentinel-data.tabulartime-series-forecasting10M<n<100M0 likes80 downloads9mo agoHugging Face27foundation-models /pi-rsi-experiment-artifacts0 likes69 downloads2mo agoHugging Face28bkai-foundation-models /vietnamese-roleplay-realm 🇻🇳 Vietnamese Role-play Realm Dataset This is a dataset of GPT-generated Vietnamese characters made to increase the ability of open-source language models to role-play. It contains 446 characters generated by GPT-3.5 Each character will have 20 topics generated by ChatGPT. And each topic will have a conversation corresponding with it In 446 characters, there are 400 general characters and 46 Vietnamese characters. To construct this dataset, we follow a four-step process:… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/vietnamese-roleplay-realm.text-generation3 likes59 downloads3y agoHugging Face29danish-foundation-models /synthetic-values-model-charter value_units.jsonl is the individual parsed values from the model charter. scenarios.jsonl is invididual hypothetical scenarios based on the values in values_units.jsonl sft_*.jsonl generated accepted responses. dpo_*.jsonl generated accepted+rejected responses. 0 likes57 downloads2mo agoHugging Face30bkai-foundation-models /vi-self-chat-sharegpt-format 🇻🇳 Vietnamese Self-Chat Dataset This dataset is designed to enhance the model's ability to engage in multi-turn conversations with humans. To construct this dataset, we follow a two-step process: Step 1: Instruction Generation We employ the methodology outlined in the Self-Instruct paper to craft a diverse set of instructions. This paper serves as a guide for aligning pretrained language models with specific instructions, providing a structured foundation for subsequent dialogue… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/vi-self-chat-sharegpt-format.text10K<n<100K13 likes49 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.