datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ko.SHP
๐ข Korean Stanford Human Preferences Dataset (Ko.SHP)
์ด ๋ฐ์ดํฐ์
์ ์์ฒด ๊ตฌ์ถํ ๋ฒ์ญ๊ธฐ๋ฅผ ํ์ฉํ์ฌ stanfordnlp/SHP ๋ฐ์ดํฐ์
์ ๋ฒ์ญํ ๊ฒ์
๋๋ค.
์๋์ ๋ด์ฉ์ ํด๋น ๋ฒ์ญ๊ธฐ๋ก README ํ์ผ์ ๋ฒ์ญํ ๊ฒ์
๋๋ค. ์ฐธ๊ณ ๋ถํ๋๋ฆฝ๋๋ค.
If you mention this dataset in a paper, please cite the paper: Understanding Dataset Difficulty with V-Usable Information (ICML 2022).
Summary
SHP๋ ์๋ฆฌ์์ ๋ฒ๋ฅ ์กฐ์ธ์ ์ด๋ฅด๊ธฐ๊น์ง 18๊ฐ์ง ๋ค๋ฅธ ์ฃผ์ ์์ญ์ ์ง๋ฌธ/์ง์นจ์ ๋ํ ์๋ต์ ๋ํ 385K ์ง๋จ ์ธ๊ฐ ์ ํธ๋ ๋ฐ์ดํฐ ์ธํธ์ด๋ค.
๊ธฐ๋ณธ ์ค์ ์ ๋ค๋ฅธ ์๋ต์ ๋ ํ ํ ์๋ต์ ์ ์ฉ์ฑ์ ๋ฐ์ ํ๊ธฐ ์ํ ๊ฒ์ด๋ฉฐ RLHF ๋ณด์ ๋ชจ๋ธ ๋ฐ NLG ํ๊ฐ ๋ชจ๋ธ (์: SteamSHP)์ ํ๋ จ ํ๋ ๋ฐโฆ See the full description on the dataset page: https://huggingface.co/datasets/nlp-with-deeplearning/ko.SHP.Ko.SlimOrca์๋ณธ ๋ฐ์ดํฐ์
: Open-Orca/SlimOrca
ko.databricks-dolly-15k์๋ณธ ๋ฐ์ดํฐ์
: databricks/databricks-dolly-15k
Corrector101zhTW
ERNIE for Chinese Spelling Correction ็น้ซไธญๆ
MacBertMaskedLM For Chinese Spelling Correction ็น้ซไธญๆ
wikipedia-zh-20230720-filtered.json ็น้ซไธญๆ
Automatic Corpus Generation-zh ็น้ซไธญๆ
้ฃไบ่ช็ถ่ช่จ่็ (Natural Language Processing, NLP) ่ธฉ็ๅ -- ๆๆฌ็ณพ้ฏ
ko.openhermes์๋ณธ ๋ฐ์ดํฐ์
: teknium/openhermes
Ko.HelpSteer์๋ณธ ๋ฐ์ดํฐ์
: nvidia/HelpSteer
Ko.WizardLM_evol_instruct_V2_196k์ด ๋ฐ์ดํฐ์
์ ์์ฒด ๊ตฌ์ถํ ๋ฒ์ญ๊ธฐ๋ก WizardLM/WizardLM_evol_instruct_V2_196k์ ๋ฒ์ญํ ๋ฐ์ดํฐ์
์
๋๋ค. ์๋ README ํ์ด์ง๋ ๋ฒ์ญ๊ธฐ๋ฅผ ํตํด ๋ฒ์ญ๋์์ต๋๋ค. ์ฐธ๊ณ ๋ถํ๋๋ฆฝ๋๋ค.
News
๐ฅ ๐ฅ ๐ฅ [08/11/2023] WizardMath ๋ชจ๋ธ์ ์ถ์ํฉ๋๋ค.
๐ฅ WizardMath-70B-V1.0 ๋ชจ๋ธ์ ChatGPT 3.5, Claude Instant 1 ๋ฐ PaLM 2 540B ๋ฅผ ํฌํจ ํ ์ฌ GSM8K์์ ์ผ๋ถ ํ์ ์์ค LLMs ๋ณด๋ค ์ฝ๊ฐ ๋ ์ฐ์ ํฉ๋๋ค.
๐ฅ ์ฐ๋ฆฌ์ WizardMath-70B-V1.0 ๋ชจ๋ธ์ SOTA ์คํ ์์ค LLM๋ณด๋ค 24.8 ํฌ์ธํธ ๋์ GSM8k Benchmarks์์ 81.6 pass@1 ์ ๋ฌ์ฑํฉ๋๋ค.
๐ฅ ์ฐ๋ฆฌ์ WizardMath-70B-V1.0 ๋ชจ๋ธ์ SOTA ์คํ ์์ค LLM๋ณด๋ค 9.2 ํฌ์ธํธ ๋์ MATH ๋ฒค์น๋งํฌ์์ 22.7 pass@1 ์ ๋ฌ์ฑํฉ๋๋ค.โฆ See the full description on the dataset page: https://huggingface.co/datasets/nlp-with-deeplearning/Ko.WizardLM_evol_instruct_V2_196k.deeplearning-tasks-v1
deeplearning-tasks-v1
36 exact tasks for the
deeplearning-env
RL environment, on the topics of Deep Learning (Goodfellow, Bengio & Courville): information
theory, backpropagation, optimisation, the linear algebra used in ML, and numerical stability.
field
meaning
task_id
dl-000 โฆ dl-035
category
information / backprop / optimisation / linalg / numerical
prompt
the question, the units, and the exact shape of the answer
api_description
the fixed network andโฆ See the full description on the dataset page: https://huggingface.co/datasets/eltociear/deeplearning-tasks-v1.ko.lima์๋ณธ ๋ฐ์ดํฐ์
: GAIR/lima
deeplearning-minimind-RL
