Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01databricks /databricks-dolly-15k Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/databricks/databricks-dolly-15k.textquestion-answering10K<n<100K1.3k likes36k downloads3y agoHugging Face02kunishou /databricks-dolly-15k-ja This dataset was created by automatically translating "databricks-dolly-15k" into Japanese.This dataset is licensed under CC-BY-SA-3.0 Last Update : 2023-05-11 databricks-dolly-15k-jahttps://github.com/kunishou/databricks-dolly-15k-jadatabricks-dolly-15khttps://github.com/databrickslabs/dolly/tree/master/data text10K<n<100K89 likes699 downloads3y agoHugging Face03llm-jp /databricks-dolly-15k-ja databricks-dolly-15k-ja This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan. This dataset is a Japanese translation of databricks-dolly-15k using DeepL. Send Questions to llm-jp(at)nii.ac.jp Model Card Authors The names are listed in alphabetical order. Hirokazu Kiyomaru, Hiroshi Matsuda, Jun Suzuki, Namgi Han, Saku Sugawara, Shota Sasaki, Shuhei Kurita, Taishi Nakamura, Takashi Kodama, Takumi… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/databricks-dolly-15k-ja.textquestion-answering10K<n<100K18 likes156 downloads3y agoHugging Face04bbz662bbz /databricks-dolly-15k-ja-gozaruThis dataset was using "kunishou/databricks-dolly-15k-ja" This dataset is licensed under CC BY SA 3.0 Last Update : 2023-05-28 databricks-dolly-15k-ja-gozaru kunishou/databricks-dolly-15k-ja https://huggingface.co/datasets/kunishou/databricks-dolly-15k-ja text10K<n<100K9 likes155 downloads3y agoHugging Face05atasoglu /databricks-dolly-15k-trThis dataset is machine-translated version of databricks-dolly-15k.jsonl into Turkish. Used googletrans==3.1.0a0 to translation. textquestion-answering10K<n<100K17 likes85 downloads3y agoHugging Face06open-llm-leaderboard /databricks__dolly-v2-7b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v2-7b Dataset automatically created during the evaluation run of model databricks/dolly-v2-7b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-7b-details.tabular10K<n<100K0 likes78 downloads2y agoHugging Face07kunishou /databricks-dolly-69k-ja-en-translationThis dataset was created by automatically translating "databricks-dolly-15k" into Japanese.This dataset contains 69K ja-en-translation task data and is licensed under CC BY SA 3.0. Last Update : 2023-04-18 databricks-dolly-15k-jahttps://github.com/kunishou/databricks-dolly-15k-jadatabricks-dolly-15khttps://github.com/databrickslabs/dolly/tree/master/data text10K<n<100K15 likes73 downloads3y agoHugging Face08takaaki-inada /databricks-dolly-15k-ja-zundamonThis dataset was based on "kunishou/databricks-dolly-15k-ja". This dataset is licensed under CC BY SA 3.0 Last Update : 2023-05-11 databricks-dolly-15k-jahttps://github.com/kunishou/databricks-dolly-15k-jadatabricks-dolly-15khttps://github.com/databrickslabs/dolly/tree/master/data text10K<n<100K12 likes62 downloads3y agoHugging Face09open-llm-leaderboard /databricks__dolly-v2-12b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v2-12b Dataset automatically created during the evaluation run of model databricks/dolly-v2-12b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-12b-details.tabular10K<n<100K0 likes59 downloads2y agoHugging Face10Elliot4AI /databricksdatabricks-dolly-15k-chinese Dataset Summary 🏡🏡🏡🏡Fine-tune Dataset:中文数据集🏡🏡🏡🏡 😀😀😀😀😀😀😀😀 这个数据集是databricks/databricks-dolly-15k的中文版本,是直接翻译过来,没有经过人为检查语法。 对databricks/databricks-dolly-15k的描述,请看他的dataset card。 😀😀😀😀😀😀😀😀 This data set is the Chinese version of databricks/databricks-dolly-15k, which is directly translated without human-checked grammar. For a description of databricks/databricks-dolly-15k, see its dataset card. textquestion-answering10K<n<100K5 likes41 downloads3y agoHugging Face11sakusakumura /databricks-dolly-15k-ja-scoredFor the English version, please click here. 概要 databricks-dolly-15k-ja-scoredはkunishou/databricks-dolly-15k-jaの派生であり、BERTScoreによって提供される翻訳品質スコアが追加されています。 このデータセットは、学術的・商業的問わずクリエイティブ・コモンズ 表示 - 継承 3.0 非移植ライセンスの条件の下で何にでも使用することができます。 翻訳の品質スコア databricks-dolly-15k-jaは、databricks-dolly-15kを機械翻訳したものです。databricks-dolly-15k-jaに含まれるデータを調べてみると、以下のような品質の悪いデータが存在することが分かりました。 inputとoutputが全く同じであるデータ outputがinstructionにコピーされているデータ 表記ゆれによって表現の一貫性が保たれていないデータ 固有名詞などの翻訳に失敗しているデータ… See the full description on the dataset page: https://huggingface.co/datasets/sakusakumura/databricks-dolly-15k-ja-scored.textquestion-answering10K<n<100K6 likes40 downloads3y agoHugging Face12open-llm-leaderboard /databricks__dolly-v2-3b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v2-3b Dataset automatically created during the evaluation run of model databricks/dolly-v2-3b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-3b-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face13open-llm-leaderboard /databricks__dolly-v1-6b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v1-6b Dataset automatically created during the evaluation run of model databricks/dolly-v1-6b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v1-6b-details.tabular10K<n<100K0 likes39 downloads2y agoHugging Face14ramachetan22 /transformed_JSON_databricks-dolly-15k.jsonl Transformed Databricks-Dolly-15k Dataset Summary The Transformed Databricks-Dolly-15k dataset is a modification of the original open-source dataset created by Databricks employees, designed to facilitate instruction-following abilities in large language models (LLMs). This version has been specifically adapted to include responses in a JSON format, enhancing its utility for tasks requiring structured output. Modifications The primary transformation applied to… See the full description on the dataset page: https://huggingface.co/datasets/ramachetan22/transformed_JSON_databricks-dolly-15k.jsonl.textquestion-answering10K<n<100K0 likes38 downloads3y agoHugging Face15bbz662bbz /databricks-dolly-15k-ja-gozarinnemonThis dataset was using "kunishou/databricks-dolly-15k-ja" This dataset is licensed under CC BY SA 3.0 Last Update : 2023-05-28 databricks-dolly-15k-ja-gozarinnemon kunishou/databricks-dolly-15k-ja https://huggingface.co/datasets/kunishou/databricks-dolly-15k-ja text10K<n<100K11 likes30 downloads3y agoHugging Face16nlpai-lab /databricks-dolly-15k-kogatedKorean translation of databricks-dolly-15k via the DeepL API Note: There are cases where multilingual data has been converted to monolingual data during batch translation to Korean using the API. Below is databricks-dolly-15k's README. Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification… See the full description on the dataset page: https://huggingface.co/datasets/nlpai-lab/databricks-dolly-15k-ko.textquestion-answering10K<n<100K31 likes29 downloads3y agoHugging Face17open-llm-leaderboard /databricks__dbrx-base-detailsgated Dataset Card for Evaluation run of databricks/dbrx-base Dataset automatically created during the evaluation run of model databricks/dbrx-base The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dbrx-base-details.tabular1K<n<10K0 likes28 downloads2y agoHugging Face18yulanfmy /databricks-qa-ja データセット概要 手動で作成したDatabricksに関する質問と回答ペアの日本語データセットです。 件数:約1,300件 情報源:Databricks HPの日本語ブログやFAQなど、データブリック社員がポストしたQitta記事 https://github.com/yulan-yan/build-your-chat-bot-JP デモに利用したデータです。 textquestion-answering1K<n<10K5 likes26 downloads3y agoHugging Face19umarzein /databricks-dolly-15k-enThis is a checkpoint of the databricks-dolly-15k dataset text10K<n<100K0 likes26 downloads3y agoHugging Face20ai-bites /databricks-miniThis is a subset of the databricks 15k dataset databricks/databricks-dolly-15k used for finetuning Google's Gemma model google/gemma-2b. This version has only those records without context to match the dataset used in the fine-tuning Keras example from Google. text10K<n<100K3 likes26 downloads3y agoHugging Face21kanhatakeyama /databricks-dolly-15k-ja-regen-nemotron databricks-dolly-15k-jaをNemotron 4-340bで再生成するテスト text10K<n<100K1 likes25 downloads2y agoHugging Face22WarriorMama777 /databricks-dolly-15k-ja_cool Overview This dataset is edited from kunishou/databricks-dolly-15k-en.It was edited so that it would be like Yuki Nagato, who appears in "The Melancholy of Haruhi Suzumiya," with an emotionless and indifferent way of speaking.In more detail, I used VS CODE etc. to replace "です、ます" and "だ、である", etc. It's a dataset for my hobby, but feel free to use it. Links… See the full description on the dataset page: https://huggingface.co/datasets/WarriorMama777/databricks-dolly-15k-ja_cool.text10K<n<100K1 likes24 downloads3y agoHugging Face23ToPo-ToPo /databricks-dolly-15k-ja-nanoja データセットの概要 以下の、主語を拙者、語尾をござるに変更したデータセットを元に、語尾を「なのじゃ」に変換したデータセットである。 license: cc-by-sa-3.0 This dataset was using "bbz662bbz/databricks-dolly-15k-ja-gozaru" This dataset is licensed under CC BY SA 3.0 bbz662bbz/databricks-dolly-15k-ja-gozaru https://huggingface.co/datasets/bbz662bbz/databricks-dolly-15k-ja-gozaru text10K<n<100K0 likes24 downloads3y agoHugging Face24nlp-with-deeplearning /ko.databricks-dolly-15k원본 데이터셋: databricks/databricks-dolly-15k textquestion-answering10K<n<100K1 likes24 downloads3y agoHugging Face25robinhad /databricks-dolly-15k-uk Summary databricks-dolly-15k-uk is an open source dataset based on databricks/databricks-dolly-15k instruction-following dataset, but machine translated using facebook/m2m100_1.2B model.Tasks covered include brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization.Expect this dataset to not be grammatically correct and having obvious pitfalls of machine translation. Original Summary # Summary `databricks-dolly-15k` is an open… See the full description on the dataset page: https://huggingface.co/datasets/robinhad/databricks-dolly-15k-uk.textquestion-answering10K<n<100K4 likes23 downloads3y agoHugging Face26higashi1 /databricks-dolly-15k-jatext10K<n<100K0 likes23 downloads3y agoHugging Face27w95 /databricks-dolly-15k-azThis dataset is a machine-translated version of databricks-dolly-15k.jsonl into Azerbaijani. Dataset size is 8k. Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization. This dataset can be used for any purpose… See the full description on the dataset page: https://huggingface.co/datasets/w95/databricks-dolly-15k-az.textquestion-answering1K<n<10K3 likes22 downloads3y agoHugging Face28Nyooti /databricks-dolly-15k Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/Nyooti/databricks-dolly-15k.textquestion-answering10K<n<100K0 likes21 downloads6mo agoHugging Face29takosama /databricks-dolly-15k-ja-google-transDolly 日本語翻訳版 このリポジトリは、Databricksが開発したdollyプロジェクトの日本語翻訳版です。 翻訳元 翻訳元のプロジェクトは以下のリンクで確認できます: Dolly(英語版) ライセンスと帰属 Copyright (2023) Databricks, Inc. このデータセットはDatabricks (https://www.databricks.com) で開発され、CC BY-SA 3.0ライセンスに基づいて使用が許可されています。 データセットの一部のカテゴリには、以下のソースからの素材が含まれており、CC BY-SA 3.0ライセンスでライセンスされています: ウィキペディア(様々なページ) - https://www.wikipedia.org/ Copyright © ウィキペディア編集者および投稿者。 この翻訳作品は、元のdollyプロジェクトがCC BY-SA 3.0で公開されているため、同じくCC BY-SA 3.0で公開しています。 詳細については、クリエイティブ・コモンズ 表示-継承 3.0ライセンスの下に提供されています。… See the full description on the dataset page: https://huggingface.co/datasets/takosama/databricks-dolly-15k-ja-google-trans.text10K<n<100K2 likes18 downloads3y agoHugging Face30QEU /databricks-dolly-16k-line_ja-2_of_4 このデータセットは、2023年に有名になったdatabrick-15kの日本語版です。 ただし、データは4分割されています。 データの内容は非常に変わっています。(半分ぐらいは、原型をとどめていません) カタカナ語にカッコ付けで英語を追記しました。 このデータセットには、QnAとして異常なレコードが見られることから修正しました。 「ゲームオブスローン」に関するトリビアなど、情報価値が低いものは削除しました。 その他、いろいろなトライアルとして情報を追加しました。 詳しい情報はこちらのブログを参考にしてください。 text1K<n<10K0 likes18 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.