datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ko-openchat-0406다음 공개된 데이터를 모두 포멧 통일 후 병합. 이후 1000개를 무작위로 추출하여 test set으로 사용
지시문 수행(Instruction-Following), 추론(Reasoning), 일반상식(Commonsense)
이 데이터들에도 수학, 코딩 데이터가 섞여있긴 합니다
FreedomIntelligence/evol-instruct-korean
heegyu/OpenOrca-gugugo-ko-len500
MarkrAI/KoCommercial-Dataset
heegyu/CoT-collection-ko
changpt/ko-lima-vicuna
maywell/koVast
dbdu/ShareGPT-74k-koHuggingFaceH4/ultrachat_200k
Open-Orca/SlimOrca-Dedup
수학, 코딩, 함수 호출 (Function Calling)
heegyu/glaive-function-calling-v2-ko… See the full description on the dataset page: https://huggingface.co/datasets/heegyu/ko-openchat-0406.lm-eval-results-openchat-openchat-3.6-8b-20240522-private
Dataset Card for Evaluation run of openchat/openchat-3.6-8b-20240522
Dataset automatically created during the evaluation run of model openchat/openchat-3.6-8b-20240522
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-openchat-openchat-3.6-8b-20240522-private.open_chat_0106_ultra_feedback_n32
Dataset Card for "open_chat_0106_ultra_feedback_n32"
More Information needed
Pretergeek__OpenChat-3.5-0106_8.99B_40Layers-Appended-details
Dataset Card for Evaluation run of Pretergeek/OpenChat-3.5-0106_8.99B_40Layers-Appended
Dataset automatically created during the evaluation run of model Pretergeek/OpenChat-3.5-0106_8.99B_40Layers-Appended
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__OpenChat-3.5-0106_8.99B_40Layers-Appended-details.openchat-spin-slimorca-iter1-dataset
Dataset Card for "openchat-spin-slimorca-iter1-dataset"
More Information needed
Pretergeek__openchat-3.5-0106_Rebased_Mistral-7B-v0.2-details
Dataset Card for Evaluation run of Pretergeek/openchat-3.5-0106_Rebased_Mistral-7B-v0.2
Dataset automatically created during the evaluation run of model Pretergeek/openchat-3.5-0106_Rebased_Mistral-7B-v0.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__openchat-3.5-0106_Rebased_Mistral-7B-v0.2-details.openchat__openchat-3.5-1210-details
Dataset Card for Evaluation run of openchat/openchat-3.5-1210
Dataset automatically created during the evaluation run of model openchat/openchat-3.5-1210
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/openchat__openchat-3.5-1210-details.ko-openchat-0404한국어 챗봇 학습을 위해, 여러 데이터를 가져와서 포멧을 통일
heegyu/glaive-function-calling-v2-ko: 15170 items
FreedomIntelligence/evol-instruct-korean: 59022 items
heegyu/PKU-SafeRLHF-ko: 135213 items
maywell/koVast: 684579 items
MarkrAI/KoCommercial-Dataset: 175454 items
HuggingFaceH4/ultrachat_200k: 207865 items
Open-Orca/SlimOrca-Dedup: 363491 items
glaiveai/glaive-code-assistant-v2: 215166 items
cogstack-opengpt-sharegpt
CogStack OpenGPT data in ShareGPT format
openchat__openchat_3.5-details
Dataset Card for Evaluation run of openchat/openchat_3.5
Dataset automatically created during the evaluation run of model openchat/openchat_3.5
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/openchat__openchat_3.5-details.openchat__openchat-3.5-0106-details
Dataset Card for Evaluation run of openchat/openchat-3.5-0106
Dataset automatically created during the evaluation run of model openchat/openchat-3.5-0106
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/openchat__openchat-3.5-0106-details.openchat__openchat-3.6-8b-20240522-details
Dataset Card for Evaluation run of openchat/openchat-3.6-8b-20240522
Dataset automatically created during the evaluation run of model openchat/openchat-3.6-8b-20240522
The dataset is composed of 43 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/openchat__openchat-3.6-8b-20240522-details.Pretergeek__OpenChat-3.5-0106_32K-PoSE-details
Dataset Card for Evaluation run of Pretergeek/OpenChat-3.5-0106_32K-PoSE
Dataset automatically created during the evaluation run of model Pretergeek/OpenChat-3.5-0106_32K-PoSE
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__OpenChat-3.5-0106_32K-PoSE-details.openchat__openchat_v3.2_super-details
Dataset Card for Evaluation run of openchat/openchat_v3.2_super
Dataset automatically created during the evaluation run of model openchat/openchat_v3.2_super
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/openchat__openchat_v3.2_super-details.beowolx__CodeNinja-1.0-OpenChat-7B-details
Dataset Card for Evaluation run of beowolx/CodeNinja-1.0-OpenChat-7B
Dataset automatically created during the evaluation run of model beowolx/CodeNinja-1.0-OpenChat-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/beowolx__CodeNinja-1.0-OpenChat-7B-details.ultrachat-sharegpt
UltraChat dataset in ShareGPT format
This is the full UltraChat dataset converted to ShareGPT format.
Pretergeek__OpenChat-3.5-0106_9.86B_44Layers-Appended-details
Dataset Card for Evaluation run of Pretergeek/OpenChat-3.5-0106_9.86B_44Layers-Appended
Dataset automatically created during the evaluation run of model Pretergeek/OpenChat-3.5-0106_9.86B_44Layers-Appended
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__OpenChat-3.5-0106_9.86B_44Layers-Appended-details.openchat__openchat_v3.2-details
Dataset Card for Evaluation run of openchat/openchat_v3.2
Dataset automatically created during the evaluation run of model openchat/openchat_v3.2
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/openchat__openchat_v3.2-details.openchatgpt-safe-r2I'm too lazy to fill in the dataset card template! Think of it like r1, but after NY - timestamp is XX-01-2023. This is not turbo at this point, it was before 26ths. This must be "alpha", I'm 99% sure.
Has same problems, additional one is missing greetings! "NDA" stuff is missing from this as well!
Pretergeek__OpenChat-3.5-0106_8.11B_36Layers-Appended-details
Dataset Card for Evaluation run of Pretergeek/OpenChat-3.5-0106_8.11B_36Layers-Appended
Dataset automatically created during the evaluation run of model Pretergeek/OpenChat-3.5-0106_8.11B_36Layers-Appended
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__OpenChat-3.5-0106_8.11B_36Layers-Appended-details.Pretergeek__OpenChat-3.5-0106_10.7B_48Layers-Appended-details
Dataset Card for Evaluation run of Pretergeek/OpenChat-3.5-0106_10.7B_48Layers-Appended
Dataset automatically created during the evaluation run of model Pretergeek/OpenChat-3.5-0106_10.7B_48Layers-Appended
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__OpenChat-3.5-0106_10.7B_48Layers-Appended-details.Pretergeek__OpenChat-3.5-0106_8.99B_40Layers-Interleaved-details
Dataset Card for Evaluation run of Pretergeek/OpenChat-3.5-0106_8.99B_40Layers-Interleaved
Dataset automatically created during the evaluation run of model Pretergeek/OpenChat-3.5-0106_8.99B_40Layers-Interleaved
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__OpenChat-3.5-0106_8.99B_40Layers-Interleaved-details.Pretergeek__OpenChat-3.5-0106_10.7B_48Layers-Interleaved-details
Dataset Card for Evaluation run of Pretergeek/OpenChat-3.5-0106_10.7B_48Layers-Interleaved
Dataset automatically created during the evaluation run of model Pretergeek/OpenChat-3.5-0106_10.7B_48Layers-Interleaved
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__OpenChat-3.5-0106_10.7B_48Layers-Interleaved-details.Pretergeek__OpenChat-3.5-0106_8.11B_36Layers-Interleaved-details
Dataset Card for Evaluation run of Pretergeek/OpenChat-3.5-0106_8.11B_36Layers-Interleaved
Dataset automatically created during the evaluation run of model Pretergeek/OpenChat-3.5-0106_8.11B_36Layers-Interleaved
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pretergeek__OpenChat-3.5-0106_8.11B_36Layers-Interleaved-details.ko-openchat-0404-test한국어 챗봇 학습을 위해, 여러 데이터를 가져와서 포멧을 통일 (각 데이터셋마다 처음 1만개씩 추출)
heegyu/glaive-function-calling-v2-ko: 15170 items
heegyu/PKU-SafeRLHF-ko: 135213 items
maywell/koVast: 684579 items
MarkrAI/KoCommercial-Dataset: 175454 items
HuggingFaceH4/ultrachat_200k: 207865 items
Open-Orca/SlimOrca-Dedup: 363491 items
glaiveai/glaive-code-assistant-v2: 215166 items
OpenChatData
OpenChatData
OpenChatData is an anonymized dataset derived from database dumps from a discontinued AI chatbot service that routed model requests through OpenRouter.
The dataset contains 20,949 chat-log records collected between February 4, 2026 and April 5, 2026, covering usage across 27 model identifiers.
Important: OpenChatData does not contain the text of user prompts or model responses. The released data consists of metadata and aggregate measurements such as token, word… See the full description on the dataset page: https://huggingface.co/datasets/gptforfree/OpenChatData.openchat_sharegpt4_dataset
openchat/openchat_sharegpt4_dataset
This is a processed version of sharegpt_clean.json from openchat/openchat_sharegpt4_dataset.
Unfortunately, there's not much information about that dataset on Huggingface and GitHub.
Changes:
Corrected turn order
Added language field for detected language
Removed duplicates
Redacted URLs, e-mails, phone numbers
Shuffled
Language code
Number of rows
en
54288
zh
8966
ko
2928
es
2193
fr
1885
ja
1402
de
823
pt
755
ru
507… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/openchat_sharegpt4_dataset.dictionary-openchat-3.5-0106To watch a video on how this dataset was created, watch the following videos:
Are words free?:
https://youtu.be/Utg_D-yQB_E?si=FKp_QZ4PbKesiDrn
Replacing Chatgpt 3.5 turbo workflows with Openchat:
https://youtu.be/DNKepnKuZns?si=bleufaiGdwGdrueK
medmcqa_mixtral_openchat_0.1
Medmcqa mixtral openchat 0.1
This dataset is a small subset of Medmcqa where asked mixtral / openchat3.5 models to answer medical questions and give some explanations.
To ensure the results are correct, we gave some useful information in the prompt to help the model to answer correct. We discarded those useful information from the question we put in this dataset.
By doing this, we can have a structured and accurate answer built by the LLM.
This dataset can be used to finetuned LLMs… See the full description on the dataset page: https://huggingface.co/datasets/guigux/medmcqa_mixtral_openchat_0.1.nectar_openchat_preprocess
