datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
orca-math-word-problems-100k-en-zh-mix100k English and Chinese mixed version of microsoft/orca-math-word-problems-200k
Vietnamese-microsoft-orca-math-word-problems-200k-gg-translatedorca-math-word-problems-200k-turkmen
Turkmen Orca Math Word Problems 200k Dataset
Overview
This dataset is a Turkmen translation of the original microsoft/orca-math-word-problems-200k dataset. The Orca Math Word Problems dataset contains 200,000 high-quality math word problems and their solutions. This Turkmen version aims to extend the accessibility of math problem-solving datasets to the Turkmen language community.
Dataset Details
Original Dataset: microsoft/orca-math-word-problems-200k… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/orca-math-word-problems-200k-turkmen.Math_Word-Problems-ShareGPTConverted, deslopped, min-hash deduplicated, rejection filtered, grammar corrected using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing
orca-math-word-problems-193k-korean-jsonl원본 데이터셋
https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k
https://huggingface.co/datasets/kuotient/orca-math-word-problems-193k-korean
Citation
@misc{mitra2024orcamath,
title={Orca-Math: Unlocking the potential of SLMs in Grade School Math},
author={Arindam Mitra and Hamed Khanpour and Corby Rosset and Ahmed Awadallah},
year={2024},
eprint={2402.14830},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
orca-math-word-problems-200k-hindi-filteredadaption-bg-math-word-problems
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-bg_math_word_problems
This dataset contains 1,319 Bulgarian language math word problems paired with step-by-step solutions that include intermediate calculations. The content covers various arithmetic scenarios such as currency conversion, cost estimation, and area calculations, formatted as prompt-completion pairs. It is designed for evaluating or training models on… See the full description on the dataset page: https://huggingface.co/datasets/Ilia-Iliev/adaption-bg-math-word-problems.math-word-problemsmicrosoft_orca-math-word-problems-200k-ShareGPT
