Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01argilla /databricks-dolly-15k-curated-en Guidelines In this dataset, you will find a collection of records that show a category, an instruction, a context and a response to that instruction. The aim of the project is to correct the instructions, intput and responses to make sure they are of the highest quality and that they match the task category that they belong to. All three texts should be clear and include real information. In addition, the response should be as complete but concise as possible. To curate the dataset… See the full description on the dataset page: https://huggingface.co/datasets/argilla/databricks-dolly-15k-curated-en.text10K<n<100K45 likes42k downloads3y agoHugging Face02databricks /databricks-dolly-15k Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/databricks/databricks-dolly-15k.textquestion-answering10K<n<100K1.3k likes36k downloads3y agoHugging Face03databricks /officeqagated OfficeQA Dataset Summary OfficeQA is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents. The benchmark consists of question–answer pairs that require reasoning over historical U.S. Treasury Bulletin documents (1939–2025), which contain dense financial tables, charts, and narrative text. OfficeQA is designed to test retrieval, tool use, and multi-step reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa.documentquestion-answeringn<1K30 likes9k downloads3mo agoHugging Face04HuggingFaceH4 /databricks_dolly_15k Dataset Card for Dolly_15K Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/databricks_dolly_15k.text10K<n<100K23 likes4.9k downloads3y agoHugging Face05databricks /officeqa-pro-v2gated OfficeQA Pro v2 Dataset Summary OfficeQA Pro v2 is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents. The benchmark consists of question–answer pairs that require reasoning over two centuries of U.S. Federal Accounts of Receipts and Expenditures reporting (1793–2024) — Combined Statements of Receipts, Outlays, and Balances of the United States Government, together with earlier… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa-pro-v2.documentquestion-answeringn<1K17 likes1.6k downloads2mo agoHugging Face06kunishou /databricks-dolly-15k-ja This dataset was created by automatically translating "databricks-dolly-15k" into Japanese.This dataset is licensed under CC-BY-SA-3.0 Last Update : 2023-05-11 databricks-dolly-15k-jahttps://github.com/kunishou/databricks-dolly-15k-jadatabricks-dolly-15khttps://github.com/databrickslabs/dolly/tree/master/data text10K<n<100K89 likes699 downloads3y agoHugging Face07argilla /databricks-dolly-15k-curated-multilingual Dataset Card for "databricks-dolly-15k-curated-multilingual" A curated and multilingual version of the Databricks Dolly instructions dataset. It includes a programmatically and manually corrected version of the original en dataset. See below. STATUS: Currently, the original Dolly v2 English version has been curated combining automatic processing and collaborative human curation using Argilla (~400 records have been manually edited and fixed). The following graph shows a summary… See the full description on the dataset page: https://huggingface.co/datasets/argilla/databricks-dolly-15k-curated-multilingual.texttext-generation10K<n<100K54 likes300 downloads3y agoHugging Face08bowang0911 /databricks-qa-ja License & Attribution MTEB-format derivative of yulanfmy/databricks-qa-ja (Japanese Databricks/Dolly-style technical QA). Query = question; corpus = answer. Licensed under CC-BY-SA-3.0 (same as source). tabulartext-retrieval1K<n<10K0 likes285 downloads4mo agoHugging Face09llm-jp /databricks-dolly-15k-ja databricks-dolly-15k-ja This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan. This dataset is a Japanese translation of databricks-dolly-15k using DeepL. Send Questions to llm-jp(at)nii.ac.jp Model Card Authors The names are listed in alphabetical order. Hirokazu Kiyomaru, Hiroshi Matsuda, Jun Suzuki, Namgi Han, Saku Sugawara, Shota Sasaki, Shuhei Kurita, Taishi Nakamura, Takashi Kodama, Takumi… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/databricks-dolly-15k-ja.textquestion-answering10K<n<100K18 likes156 downloads3y agoHugging Face10bbz662bbz /databricks-dolly-15k-ja-gozaruThis dataset was using "kunishou/databricks-dolly-15k-ja" This dataset is licensed under CC BY SA 3.0 Last Update : 2023-05-28 databricks-dolly-15k-ja-gozaru kunishou/databricks-dolly-15k-ja https://huggingface.co/datasets/kunishou/databricks-dolly-15k-ja text10K<n<100K9 likes155 downloads3y agoHugging Face11open-llm-leaderboard-old /details_databricks__dbrx-instruct0 likes131 downloads2y agoHugging Face12atasoglu /databricks-dolly-15k-trThis dataset is machine-translated version of databricks-dolly-15k.jsonl into Turkish. Used googletrans==3.1.0a0 to translation. textquestion-answering10K<n<100K17 likes85 downloads3y agoHugging Face13open-llm-leaderboard-old /details_databricks__dbrx-base Dataset Card for Evaluation run of databricks/dbrx-base Dataset automatically created during the evaluation run of model databricks/dbrx-base on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_databricks__dbrx-base.0 likes81 downloads3y agoHugging Face14open-llm-leaderboard /databricks__dolly-v2-7b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v2-7b Dataset automatically created during the evaluation run of model databricks/dolly-v2-7b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-7b-details.tabular10K<n<100K0 likes78 downloads2y agoHugging Face15JLongshore /databricks-cost-leak-hunter Databricks Cost Leak Hunter — Agent Skill Hunt down Databricks cost leaks — wasted DBUs, idle clusters, oversized SQL warehouses, and untagged runaway spend — and produce a FinOps cost report your CFO can read. A self-contained Agent Skill for Claude Code, OpenAI Codex, Cursor, and Gemini CLI. This is the full production skill: SKILL.md plus reference docs (cost-leak taxonomy, system-tables setup, DLT tier tradeoffs, CFO output format), helper scripts (spend-baseline SQL… See the full description on the dataset page: https://huggingface.co/datasets/JLongshore/databricks-cost-leak-hunter.0 likes77 downloads3mo agoHugging Face16kunishou /databricks-dolly-69k-ja-en-translationThis dataset was created by automatically translating "databricks-dolly-15k" into Japanese.This dataset contains 69K ja-en-translation task data and is licensed under CC BY SA 3.0. Last Update : 2023-04-18 databricks-dolly-15k-jahttps://github.com/kunishou/databricks-dolly-15k-jadatabricks-dolly-15khttps://github.com/databrickslabs/dolly/tree/master/data text10K<n<100K15 likes73 downloads3y agoHugging Face17open-llm-leaderboard-old /details_databricks__dolly-v2-7b Dataset Card for Evaluation run of databricks/dolly-v2-7b Dataset Summary Dataset automatically created during the evaluation run of model databricks/dolly-v2-7b on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_databricks__dolly-v2-7b.0 likes71 downloads3y agoHugging Face18open-llm-leaderboard-old /details_databricks__dolly-v2-12b Dataset Card for Evaluation run of databricks/dolly-v2-12b Dataset Summary Dataset automatically created during the evaluation run of model databricks/dolly-v2-12b on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_databricks__dolly-v2-12b.0 likes70 downloads3y agoHugging Face19lilacai /lilac-databricks-dolly-15k-curated-en lilac/databricks-dolly-15k-curated-en This dataset is a Lilac processed dataset. Original dataset: https://huggingface.co/datasets/argilla/databricks-dolly-15k-curated-en To download the dataset to a local directory: lilac download lilacai/lilac-databricks-dolly-15k-curated-en or from python with: ll.download("lilacai/lilac-databricks-dolly-15k-curated-en") 0 likes68 downloads3y agoHugging Face20nahsa /databricks_dolly_15k Dataset Card for Dolly_15K Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/nahsa/databricks_dolly_15k.text10K<n<100K0 likes66 downloads22d agoHugging Face21takaaki-inada /databricks-dolly-15k-ja-zundamonThis dataset was based on "kunishou/databricks-dolly-15k-ja". This dataset is licensed under CC BY SA 3.0 Last Update : 2023-05-11 databricks-dolly-15k-jahttps://github.com/kunishou/databricks-dolly-15k-jadatabricks-dolly-15khttps://github.com/databrickslabs/dolly/tree/master/data text10K<n<100K12 likes62 downloads3y agoHugging Face22open-llm-leaderboard-old /details_databricks__dolly-v2-3b Dataset Card for Evaluation run of databricks/dolly-v2-3b Dataset Summary Dataset automatically created during the evaluation run of model databricks/dolly-v2-3b on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_databricks__dolly-v2-3b.0 likes59 downloads3y agoHugging Face23open-llm-leaderboard /databricks__dolly-v2-12b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v2-12b Dataset automatically created during the evaluation run of model databricks/dolly-v2-12b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-12b-details.tabular10K<n<100K0 likes59 downloads2y agoHugging Face24manishiitg /databricks-databricks-dolly-15ktext10K<n<100K0 likes55 downloads3y agoHugging Face25aisquared /databricks-dolly-15k databricks-dolly-15k This dataset was not originally created by AI Squared. This dataset was curated and created by Databricks. The below text comes from the original release of the dataset's README file in GitHub (available at https://github.com/databrickslabs/dolly/tree/master/data): Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in… See the full description on the dataset page: https://huggingface.co/datasets/aisquared/databricks-dolly-15k.text10K<n<100K7 likes47 downloads3y agoHugging Face26dariolopez /Llama-2-databricks-dolly-oasst1-es-lower-1024-tokens Llama-2-databricks-dolly-oasst1-es-lower-1024-tokens Union of https://huggingface.co/datasets/dariolopez/Llama-2-databricks-dolly-es and https://huggingface.co/datasets/dariolopez/Llama-2-oasst1-es Filtering of texts with less than 1024 tokens. text10K<n<100K0 likes45 downloads3y agoHugging Face27GenAIDevTOProd /databricks-dolly15k-semantic-complexity Databricks - Dolly 15k – Enriched Variant (Instruction-Tuned with Semantic and Complexity Augmentation) Overview This dataset is a semantically enriched and complexity-aware extension of the original Databricks Dolly 15k, purpose-built for evaluating and training instruction-following models. Each sample is augmented with additional signals to enable more nuanced filtering, curriculum learning, and benchmark development across diverse NLP tasks. Dataset Format Each… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/databricks-dolly15k-semantic-complexity.tabular10K<n<100K1 likes43 downloads1y agoHugging Face28aaqibsaeed /databricks-dolly-15k-urThis dataset was created by translating "databricks-dolly-15k.jsonl" into Urdu. It is licensed under CC BY 3.0. .اس ڈیٹا سیٹ کو "ڈیٹابرکس-ڈولی" کو اردو میں ترجمہ کرکے تیار کیا گیا تھا databricks-dolly-15k https://github.com/databrickslabs/dolly/tree/master/data text10K<n<100K2 likes42 downloads3y agoHugging Face29Elliot4AI /databricksdatabricks-dolly-15k-chinese Dataset Summary 🏡🏡🏡🏡Fine-tune Dataset:中文数据集🏡🏡🏡🏡 😀😀😀😀😀😀😀😀 这个数据集是databricks/databricks-dolly-15k的中文版本,是直接翻译过来,没有经过人为检查语法。 对databricks/databricks-dolly-15k的描述,请看他的dataset card。 😀😀😀😀😀😀😀😀 This data set is the Chinese version of databricks/databricks-dolly-15k, which is directly translated without human-checked grammar. For a description of databricks/databricks-dolly-15k, see its dataset card. textquestion-answering10K<n<100K5 likes41 downloads3y agoHugging Face30mguechaoui /databricks-dolly-15k-darijatext10K<n<100K0 likes41 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.