Team Ai
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01databricks /officeqagated OfficeQA Dataset Summary OfficeQA is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents. The benchmark consists of question–answer pairs that require reasoning over historical U.S. Treasury Bulletin documents (1939–2025), which contain dense financial tables, charts, and narrative text. OfficeQA is designed to test retrieval, tool use, and multi-step reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa.documentquestion-answeringn<1K30 likes9k downloads3mo agoHugging Face02databricks /officeqa-pro-v2gated OfficeQA Pro v2 Dataset Summary OfficeQA Pro v2 is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents. The benchmark consists of question–answer pairs that require reasoning over two centuries of U.S. Federal Accounts of Receipts and Expenditures reporting (1793–2024) — Combined Statements of Receipts, Outlays, and Balances of the United States Government, together with earlier… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa-pro-v2.documentquestion-answeringn<1K17 likes1.6k downloads2mo agoHugging Face03argilla /databricks-dolly-15k-curated-multilingual Dataset Card for "databricks-dolly-15k-curated-multilingual" A curated and multilingual version of the Databricks Dolly instructions dataset. It includes a programmatically and manually corrected version of the original en dataset. See below. STATUS: Currently, the original Dolly v2 English version has been curated combining automatic processing and collaborative human curation using Argilla (~400 records have been manually edited and fixed). The following graph shows a summary… See the full description on the dataset page: https://huggingface.co/datasets/argilla/databricks-dolly-15k-curated-multilingual.texttext-generation10K<n<100K54 likes300 downloads3y agoHugging Face04Elliot4AI /databricksdatabricks-dolly-15k-chinese Dataset Summary 🏡🏡🏡🏡Fine-tune Dataset:中文数据集🏡🏡🏡🏡 😀😀😀😀😀😀😀😀 这个数据集是databricks/databricks-dolly-15k的中文版本,是直接翻译过来,没有经过人为检查语法。 对databricks/databricks-dolly-15k的描述,请看他的dataset card。 😀😀😀😀😀😀😀😀 This data set is the Chinese version of databricks/databricks-dolly-15k, which is directly translated without human-checked grammar. For a description of databricks/databricks-dolly-15k, see its dataset card. textquestion-answering10K<n<100K5 likes41 downloads3y agoHugging Face05Felladrin /ChatML-databricks-dolly-15kdatabricks/databricks-dolly-15k in ChatML format. Python code used for conversion: from datasets import load_dataset import pandas from transformers import AutoTokenizer tokenizer = AutoTokenizer.from_pretrained( pretrained_model_name_or_path="Felladrin/Llama-160M-Chat-v1" ) dataset = load_dataset("databricks/databricks-dolly-15k", split="train") def format(columns): instruction = columns["instruction"].strip() context = columns["context"].strip() response =… See the full description on the dataset page: https://huggingface.co/datasets/Felladrin/ChatML-databricks-dolly-15k.textquestion-answering10K<n<100K1 likes24 downloads3y agoHugging Face06Sadanto3933 /databricks-sft-15ktextquestion-answering10K<n<100K3 likes18 downloads2y agoHugging Face07Alberto1231 /databricks_dolly_15k Databricks Dolly task samples Standalone task subsets derived from databricks/databricks-dolly-15k at revision bdd27f4d94b9c1f951818a7da7fd7aeea5dbff1a: general_qa (source category: general_qa) open_qa (source category: open_qa) closed_qa (source category: closed_qa) brainstorm (source category: brainstorming) classify (source category: classification) extract_information (source category: information_extraction) summarize (source category: summarization) creative_writing… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/databricks_dolly_15k.tabulartext-generationn<1K0 likes17 downloads2mo agoHugging Face08davidquicast /databricks-dolly-15k-esTranslated with googletrans==3.1.0a0 from original dataset *part of the data (up to 600) was lost during the translation license: apache-2.0 texttext-generation10K<n<100K3 likes16 downloads3y agoHugging Face09Sharathhebbar24 /databricks-dolly-15k Databricks-dolly This is a cleansed version of databricks/databricks-dolly-15k Usage from datasets import load_dataset dataset = load_dataset("Sharathhebbar24/databricks-dolly-15k", split="train") texttext-generation10K<n<100K0 likes12 downloads3y agoHugging Face10GreenNode /SFT_databricks_dolly_15k Preparing Your Dataset Once you’ve decided that fine-tuning is the best approach—after optimizing your prompt as much as possible and identifying remaining model issues—you’ll need to prepare training data. Start by creating a diverse set of example conversations that mirror those the model will handle during production. Each example should follow this structure below, consisting of a list of messages. Each message must include a role, content, and an optional name. Make sure some… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/SFT_databricks_dolly_15k.texttext-generation10K<n<100K0 likes10 downloads2y agoHugging Face11Vishaltiwari2019 /textGen-databricks-dollytexttext-generation10K<n<100K4 likes8 downloads3y agoHugging Face12MagicaNeko /databricks-dolly-1k Databricks Dolly 1k 1092 instruction examples taken from the original databricks/databricks-dolly-15k. Filtered to open/closed/general QA category Ready to plug straight into SFTTrainer, Unsloth, Llama-factory etc Example ### Instruction: When did Virgin Australia start operating? ### Context: Virgin Australia, the trading name of Virgin Australia Airlines Pty Ltd ... ### Response: Virgin Australia commenced services on 31 August 2000 as Virgin Blue, with two aircraft… See the full description on the dataset page: https://huggingface.co/datasets/MagicaNeko/databricks-dolly-1k.texttext-generation1K<n<10K0 likes8 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.