datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
databricks-dolly-15k-curated-en
Guidelines
In this dataset, you will find a collection of records that show a category, an instruction, a context and a response to that instruction. The aim of the project is to correct the instructions, intput and responses to make sure they are of the highest quality and that they match the task category that they belong to. All three texts should be clear and include real information. In addition, the response should be as complete but concise as possible.
To curate the dataset… See the full description on the dataset page: https://huggingface.co/datasets/argilla/databricks-dolly-15k-curated-en.databricks-dolly-15k
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several
of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification,
closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/databricks/databricks-dolly-15k.officeqa
OfficeQA
Dataset Summary
OfficeQA is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents.
The benchmark consists of question–answer pairs that require reasoning over historical U.S. Treasury Bulletin documents (1939–2025), which contain dense financial tables, charts, and narrative text. OfficeQA is designed to test retrieval, tool use, and multi-step reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa.databricks_dolly_15k
Dataset Card for Dolly_15K
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/databricks_dolly_15k.officeqa-pro-v2
OfficeQA Pro v2
Dataset Summary
OfficeQA Pro v2 is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents.
The benchmark consists of question–answer pairs that require reasoning over two centuries of U.S. Federal Accounts of Receipts and Expenditures reporting (1793–2024) — Combined Statements of Receipts, Outlays, and Balances of the United States Government, together with earlier… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa-pro-v2.databricks-dolly-15k-ja
This dataset was created by automatically translating "databricks-dolly-15k" into Japanese.This dataset is licensed under CC-BY-SA-3.0
Last Update : 2023-05-11
databricks-dolly-15k-jahttps://github.com/kunishou/databricks-dolly-15k-jadatabricks-dolly-15khttps://github.com/databrickslabs/dolly/tree/master/data
databricks-dolly-15k-curated-multilingual
Dataset Card for "databricks-dolly-15k-curated-multilingual"
A curated and multilingual version of the Databricks Dolly instructions dataset. It includes a programmatically and manually corrected version of the original en dataset. See below.
STATUS:
Currently, the original Dolly v2 English version has been curated combining automatic processing and collaborative human curation using Argilla (~400 records have been manually edited and fixed). The following graph shows a summary… See the full description on the dataset page: https://huggingface.co/datasets/argilla/databricks-dolly-15k-curated-multilingual.databricks-qa-ja
License & Attribution
MTEB-format derivative of yulanfmy/databricks-qa-ja (Japanese Databricks/Dolly-style technical QA). Query = question; corpus = answer. Licensed under CC-BY-SA-3.0 (same as source).
databricks-dolly-15k-ja
databricks-dolly-15k-ja
This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan.
This dataset is a Japanese translation of databricks-dolly-15k using DeepL.
Send Questions to
llm-jp(at)nii.ac.jp
Model Card Authors
The names are listed in alphabetical order.
Hirokazu Kiyomaru, Hiroshi Matsuda, Jun Suzuki, Namgi Han, Saku Sugawara, Shota Sasaki, Shuhei Kurita, Taishi Nakamura, Takashi Kodama, Takumi… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/databricks-dolly-15k-ja.databricks-dolly-15k-ja-gozaruThis dataset was using "kunishou/databricks-dolly-15k-ja"
This dataset is licensed under CC BY SA 3.0
Last Update : 2023-05-28
databricks-dolly-15k-ja-gozaru
kunishou/databricks-dolly-15k-ja
https://huggingface.co/datasets/kunishou/databricks-dolly-15k-ja
databricks-dolly-15k-trThis dataset is machine-translated version of databricks-dolly-15k.jsonl into Turkish.
Used googletrans==3.1.0a0 to translation.
databricks__dolly-v2-7b-details
Dataset Card for Evaluation run of databricks/dolly-v2-7b
Dataset automatically created during the evaluation run of model databricks/dolly-v2-7b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-7b-details.databricks-dolly-69k-ja-en-translationThis dataset was created by automatically translating "databricks-dolly-15k" into Japanese.This dataset contains 69K ja-en-translation task data and is licensed under CC BY SA 3.0.
Last Update : 2023-04-18
databricks-dolly-15k-jahttps://github.com/kunishou/databricks-dolly-15k-jadatabricks-dolly-15khttps://github.com/databrickslabs/dolly/tree/master/data
databricks_dolly_15k
Dataset Card for Dolly_15K
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/nahsa/databricks_dolly_15k.databricks-dolly-15k-ja-zundamonThis dataset was based on "kunishou/databricks-dolly-15k-ja".
This dataset is licensed under CC BY SA 3.0
Last Update : 2023-05-11
databricks-dolly-15k-jahttps://github.com/kunishou/databricks-dolly-15k-jadatabricks-dolly-15khttps://github.com/databrickslabs/dolly/tree/master/data
databricks__dolly-v2-12b-details
Dataset Card for Evaluation run of databricks/dolly-v2-12b
Dataset automatically created during the evaluation run of model databricks/dolly-v2-12b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-12b-details.databricks-databricks-dolly-15kdatabricks-dolly-15k
databricks-dolly-15k
This dataset was not originally created by AI Squared. This dataset was curated and created by Databricks.
The below text comes from the original release of the dataset's README file in GitHub (available at https://github.com/databrickslabs/dolly/tree/master/data):
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in… See the full description on the dataset page: https://huggingface.co/datasets/aisquared/databricks-dolly-15k.Llama-2-databricks-dolly-oasst1-es-lower-1024-tokens
Llama-2-databricks-dolly-oasst1-es-lower-1024-tokens
Union of https://huggingface.co/datasets/dariolopez/Llama-2-databricks-dolly-es and https://huggingface.co/datasets/dariolopez/Llama-2-oasst1-es
Filtering of texts with less than 1024 tokens.
databricks-dolly15k-semantic-complexity
Databricks - Dolly 15k – Enriched Variant (Instruction-Tuned with Semantic and Complexity Augmentation)
Overview
This dataset is a semantically enriched and complexity-aware extension of the original Databricks Dolly 15k, purpose-built for evaluating and training instruction-following models. Each sample is augmented with additional signals to enable more nuanced filtering, curriculum learning, and benchmark development across diverse NLP tasks.
Dataset Format
Each… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/databricks-dolly15k-semantic-complexity.databricks-dolly-15k-urThis dataset was created by translating "databricks-dolly-15k.jsonl" into Urdu. It is licensed under CC BY 3.0.
.اس ڈیٹا سیٹ کو "ڈیٹابرکس-ڈولی" کو اردو میں ترجمہ کرکے تیار کیا گیا تھا
databricks-dolly-15k https://github.com/databrickslabs/dolly/tree/master/data
databricksdatabricks-dolly-15k-chinese
Dataset Summary
🏡🏡🏡🏡Fine-tune Dataset:中文数据集🏡🏡🏡🏡
😀😀😀😀😀😀😀😀 这个数据集是databricks/databricks-dolly-15k的中文版本,是直接翻译过来,没有经过人为检查语法。 对databricks/databricks-dolly-15k的描述,请看他的dataset card。
😀😀😀😀😀😀😀😀 This data set is the Chinese version of databricks/databricks-dolly-15k, which is directly translated without human-checked grammar. For a description of databricks/databricks-dolly-15k, see its dataset card.
databricks-dolly-15k-darijadatabricks-dolly-15k-ja-scoredFor the English version, please click here.
概要
databricks-dolly-15k-ja-scoredはkunishou/databricks-dolly-15k-jaの派生であり、BERTScoreによって提供される翻訳品質スコアが追加されています。
このデータセットは、学術的・商業的問わずクリエイティブ・コモンズ 表示 - 継承 3.0 非移植ライセンスの条件の下で何にでも使用することができます。
翻訳の品質スコア
databricks-dolly-15k-jaは、databricks-dolly-15kを機械翻訳したものです。databricks-dolly-15k-jaに含まれるデータを調べてみると、以下のような品質の悪いデータが存在することが分かりました。
inputとoutputが全く同じであるデータ
outputがinstructionにコピーされているデータ
表記ゆれによって表現の一貫性が保たれていないデータ
固有名詞などの翻訳に失敗しているデータ… See the full description on the dataset page: https://huggingface.co/datasets/sakusakumura/databricks-dolly-15k-ja-scored.databricks__dolly-v2-3b-details
Dataset Card for Evaluation run of databricks/dolly-v2-3b
Dataset automatically created during the evaluation run of model databricks/dolly-v2-3b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-3b-details.databricks_dolly15k_endatabricks__dolly-v1-6b-details
Dataset Card for Evaluation run of databricks/dolly-v1-6b
Dataset automatically created during the evaluation run of model databricks/dolly-v1-6b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v1-6b-details.transformed_JSON_databricks-dolly-15k.jsonl
Transformed Databricks-Dolly-15k Dataset
Summary
The Transformed Databricks-Dolly-15k dataset is a modification of the original open-source dataset created by Databricks employees, designed to facilitate instruction-following abilities in large language models (LLMs). This version has been specifically adapted to include responses in a JSON format, enhancing its utility for tasks requiring structured output.
Modifications
The primary transformation applied to… See the full description on the dataset page: https://huggingface.co/datasets/ramachetan22/transformed_JSON_databricks-dolly-15k.jsonl.databricks-thinking
databricks-thinking
Created by extracing the questions from the [databricks-dolly] dataset and using Qwen3-14b to synthetically generate reasoning traces and answers.
The whole process took 1 day 23 hours 18 minuntes and 8 seconds
Why we created this dataset
The vast majority of publicly available datasets comes from large models such as DeepSeek R1.
The issue with using these large models are obvious: the reasoning traces are extremely long, often longer than the actual… See the full description on the dataset page: https://huggingface.co/datasets/chimbiwide/databricks-thinking.databricks-dolly-15k-ja-alpaca-formatThis dataset is a translation of "databricks-dolly-15k-ja", which was created by automatically translating "databricks-dolly-15k" into Japanese, into input and output formats.
This dataset is licensed under CC BY SA 3.0
Last Update : 2023-06-15
databricks-dolly-15k-ja
https://github.com/kunishou/databricks-dolly-15k-ja
databricks-dolly-15k
https://github.com/databrickslabs/dolly/tree/master/data
