datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ko.SHP
π’ Korean Stanford Human Preferences Dataset (Ko.SHP)
μ΄ λ°μ΄ν°μ
μ μ체 ꡬμΆν λ²μκΈ°λ₯Ό νμ©νμ¬ stanfordnlp/SHP λ°μ΄ν°μ
μ λ²μν κ²μ
λλ€.
μλμ λ΄μ©μ ν΄λΉ λ²μκΈ°λ‘ README νμΌμ λ²μν κ²μ
λλ€. μ°Έκ³ λΆνλ립λλ€.
If you mention this dataset in a paper, please cite the paper: Understanding Dataset Difficulty with V-Usable Information (ICML 2022).
Summary
SHPλ μ리μμ λ²λ₯ μ‘°μΈμ μ΄λ₯΄κΈ°κΉμ§ 18κ°μ§ λ€λ₯Έ μ£Όμ μμμ μ§λ¬Έ/μ§μΉ¨μ λν μλ΅μ λν 385K μ§λ¨ μΈκ° μ νΈλ λ°μ΄ν° μΈνΈμ΄λ€.
κΈ°λ³Έ μ€μ μ λ€λ₯Έ μλ΅μ λ ν ν μλ΅μ μ μ©μ±μ λ°μ νκΈ° μν κ²μ΄λ©° RLHF 보μ λͺ¨λΈ λ° NLG νκ° λͺ¨λΈ (μ: SteamSHP)μ νλ ¨ νλ λ°β¦ See the full description on the dataset page: https://huggingface.co/datasets/nlp-with-deeplearning/ko.SHP.Ko.WizardLM_evol_instruct_V2_196kμ΄ λ°μ΄ν°μ
μ μ체 ꡬμΆν λ²μκΈ°λ‘ WizardLM/WizardLM_evol_instruct_V2_196kμ λ²μν λ°μ΄ν°μ
μ
λλ€. μλ README νμ΄μ§λ λ²μκΈ°λ₯Ό ν΅ν΄ λ²μλμμ΅λλ€. μ°Έκ³ λΆνλ립λλ€.
News
π₯ π₯ π₯ [08/11/2023] WizardMath λͺ¨λΈμ μΆμν©λλ€.
π₯ WizardMath-70B-V1.0 λͺ¨λΈμ ChatGPT 3.5, Claude Instant 1 λ° PaLM 2 540B λ₯Ό ν¬ν¨ ν μ¬ GSM8Kμμ μΌλΆ νμ μμ€ LLMs λ³΄λ€ μ½κ° λ μ°μ ν©λλ€.
π₯ μ°λ¦¬μ WizardMath-70B-V1.0 λͺ¨λΈμ SOTA μ€ν μμ€ LLMλ³΄λ€ 24.8 ν¬μΈνΈ λμ GSM8k Benchmarksμμ 81.6 pass@1 μ λ¬μ±ν©λλ€.
π₯ μ°λ¦¬μ WizardMath-70B-V1.0 λͺ¨λΈμ SOTA μ€ν μμ€ LLMλ³΄λ€ 9.2 ν¬μΈνΈ λμ MATH λ²€μΉλ§ν¬μμ 22.7 pass@1 μ λ¬μ±ν©λλ€.β¦ See the full description on the dataset page: https://huggingface.co/datasets/nlp-with-deeplearning/Ko.WizardLM_evol_instruct_V2_196k.ko.databricks-dolly-15kμλ³Έ λ°μ΄ν°μ
: databricks/databricks-dolly-15k
Ko.SlimOrcaμλ³Έ λ°μ΄ν°μ
: Open-Orca/SlimOrca
