openchat
Datasets
All datasets matching “openchat”openchat_sharegpt4_datasetThis repository contains cleaned and filtered ShareGPT GPT-4 data used to train OpenChat. Details can be found in the OpenChat repository.
3_4_fusechat_v1_openchat-3.5_mixtral-8x7b-instruct-v0.1_solar-10.7b-instruct-v1.0_representationko-openchat-0406다음 공개된 데이터를 모두 포멧 통일 후 병합. 이후 1000개를 무작위로 추출하여 test set으로 사용
지시문 수행(Instruction-Following), 추론(Reasoning), 일반상식(Commonsense)
이 데이터들에도 수학, 코딩 데이터가 섞여있긴 합니다
FreedomIntelligence/evol-instruct-korean
heegyu/OpenOrca-gugugo-ko-len500
MarkrAI/KoCommercial-Dataset
heegyu/CoT-collection-ko
changpt/ko-lima-vicuna
maywell/koVast
dbdu/ShareGPT-74k-koHuggingFaceH4/ultrachat_200k
Open-Orca/SlimOrca-Dedup
수학, 코딩, 함수 호출 (Function Calling)
heegyu/glaive-function-calling-v2-ko… See the full description on the dataset page: https://huggingface.co/datasets/heegyu/ko-openchat-0406.openchat_sharegpt_v3ShareGPT dataset for training OpenChat V3 series. See OpenChat repository for instructions.
Contents:
sharegpt_clean.json: ShareGPT dataset in original format, converted to Markdown, and with model labels.
sharegpt_gpt4.json: All instances in sharegpt_clean.json with model == "Model: GPT-4".
*.parquet: Pre-tokenized dataset for training specified version of OpenChat.
Note: The dataset is NOT currently compatible with HF dataset loader.
Licensed under MIT.
lm-eval-results-openchat-openchat-3.6-8b-20240522-private
Dataset Card for Evaluation run of openchat/openchat-3.6-8b-20240522
Dataset automatically created during the evaluation run of model openchat/openchat-3.6-8b-20240522
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-openchat-openchat-3.6-8b-20240522-private.details_openchat__openchat-3.5-0106
Dataset Card for Evaluation run of openchat/openchat-3.5-0106
Dataset automatically created during the evaluation run of model openchat/openchat-3.5-0106 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_openchat__openchat-3.5-0106.
