datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ko.SHP
๐ข Korean Stanford Human Preferences Dataset (Ko.SHP)
์ด ๋ฐ์ดํฐ์
์ ์์ฒด ๊ตฌ์ถํ ๋ฒ์ญ๊ธฐ๋ฅผ ํ์ฉํ์ฌ stanfordnlp/SHP ๋ฐ์ดํฐ์
์ ๋ฒ์ญํ ๊ฒ์
๋๋ค.
์๋์ ๋ด์ฉ์ ํด๋น ๋ฒ์ญ๊ธฐ๋ก README ํ์ผ์ ๋ฒ์ญํ ๊ฒ์
๋๋ค. ์ฐธ๊ณ ๋ถํ๋๋ฆฝ๋๋ค.
If you mention this dataset in a paper, please cite the paper: Understanding Dataset Difficulty with V-Usable Information (ICML 2022).
Summary
SHP๋ ์๋ฆฌ์์ ๋ฒ๋ฅ ์กฐ์ธ์ ์ด๋ฅด๊ธฐ๊น์ง 18๊ฐ์ง ๋ค๋ฅธ ์ฃผ์ ์์ญ์ ์ง๋ฌธ/์ง์นจ์ ๋ํ ์๋ต์ ๋ํ 385K ์ง๋จ ์ธ๊ฐ ์ ํธ๋ ๋ฐ์ดํฐ ์ธํธ์ด๋ค.
๊ธฐ๋ณธ ์ค์ ์ ๋ค๋ฅธ ์๋ต์ ๋ ํ ํ ์๋ต์ ์ ์ฉ์ฑ์ ๋ฐ์ ํ๊ธฐ ์ํ ๊ฒ์ด๋ฉฐ RLHF ๋ณด์ ๋ชจ๋ธ ๋ฐ NLG ํ๊ฐ ๋ชจ๋ธ (์: SteamSHP)์ ํ๋ จ ํ๋ ๋ฐโฆ See the full description on the dataset page: https://huggingface.co/datasets/nlp-with-deeplearning/ko.SHP.Corrector101zhTW
ERNIE for Chinese Spelling Correction ็น้ซไธญๆ
MacBertMaskedLM For Chinese Spelling Correction ็น้ซไธญๆ
wikipedia-zh-20230720-filtered.json ็น้ซไธญๆ
Automatic Corpus Generation-zh ็น้ซไธญๆ
้ฃไบ่ช็ถ่ช่จ่็ (Natural Language Processing, NLP) ่ธฉ็ๅ -- ๆๆฌ็ณพ้ฏ
repro-possibilistic-predictive-uncertainty-for-deep-learning-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Ko.WizardLM_evol_instruct_V2_196k์ด ๋ฐ์ดํฐ์
์ ์์ฒด ๊ตฌ์ถํ ๋ฒ์ญ๊ธฐ๋ก WizardLM/WizardLM_evol_instruct_V2_196k์ ๋ฒ์ญํ ๋ฐ์ดํฐ์
์
๋๋ค. ์๋ README ํ์ด์ง๋ ๋ฒ์ญ๊ธฐ๋ฅผ ํตํด ๋ฒ์ญ๋์์ต๋๋ค. ์ฐธ๊ณ ๋ถํ๋๋ฆฝ๋๋ค.
News
๐ฅ ๐ฅ ๐ฅ [08/11/2023] WizardMath ๋ชจ๋ธ์ ์ถ์ํฉ๋๋ค.
๐ฅ WizardMath-70B-V1.0 ๋ชจ๋ธ์ ChatGPT 3.5, Claude Instant 1 ๋ฐ PaLM 2 540B ๋ฅผ ํฌํจ ํ ์ฌ GSM8K์์ ์ผ๋ถ ํ์ ์์ค LLMs ๋ณด๋ค ์ฝ๊ฐ ๋ ์ฐ์ ํฉ๋๋ค.
๐ฅ ์ฐ๋ฆฌ์ WizardMath-70B-V1.0 ๋ชจ๋ธ์ SOTA ์คํ ์์ค LLM๋ณด๋ค 24.8 ํฌ์ธํธ ๋์ GSM8k Benchmarks์์ 81.6 pass@1 ์ ๋ฌ์ฑํฉ๋๋ค.
๐ฅ ์ฐ๋ฆฌ์ WizardMath-70B-V1.0 ๋ชจ๋ธ์ SOTA ์คํ ์์ค LLM๋ณด๋ค 9.2 ํฌ์ธํธ ๋์ MATH ๋ฒค์น๋งํฌ์์ 22.7 pass@1 ์ ๋ฌ์ฑํฉ๋๋ค.โฆ See the full description on the dataset page: https://huggingface.co/datasets/nlp-with-deeplearning/Ko.WizardLM_evol_instruct_V2_196k.ko.databricks-dolly-15k์๋ณธ ๋ฐ์ดํฐ์
: databricks/databricks-dolly-15k
ko.openhermes์๋ณธ ๋ฐ์ดํฐ์
: teknium/openhermes
Ko.SlimOrca์๋ณธ ๋ฐ์ดํฐ์
: Open-Orca/SlimOrca
Ko.HelpSteer์๋ณธ ๋ฐ์ดํฐ์
: nvidia/HelpSteer
deeplearning-tasks-v1
deeplearning-tasks-v1
36 exact tasks for the
deeplearning-env
RL environment, on the topics of Deep Learning (Goodfellow, Bengio & Courville): information
theory, backpropagation, optimisation, the linear algebra used in ML, and numerical stability.
field
meaning
task_id
dl-000 โฆ dl-035
category
information / backprop / optimisation / linalg / numerical
prompt
the question, the units, and the exact shape of the answer
api_description
the fixed network andโฆ See the full description on the dataset page: https://huggingface.co/datasets/eltociear/deeplearning-tasks-v1.ko.lima์๋ณธ ๋ฐ์ดํฐ์
: GAIR/lima
deeplearning-minimind-RLim-map-dataset-test-deep-learning
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [Moreโฆ See the full description on the dataset page: https://huggingface.co/datasets/pioivenium/im-map-dataset-test-deep-learning.
