nassimjp/problem_solving-reasoning-pashto-plus
Problem Solving & Reasoning Multilingual Dataset This repository contains a specialized dataset focused on logical reasoning, problem-solving, and step-by-step cognitive workflows across multiple regional languages: Pashto (ps), Arabic (ar), Farsi (fa), Sindhi (sd), and Urdu (ur). Dataset Overview Languages: Pashto (پښتو), Arabic (العربية), Farsi (فارسی), Sindhi (سنڌي), Urdu (اردو) Domain: Logical Reasoning, Problem Solving, Cognitive SFT License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/problem_solving-reasoning-pashto-plus.
Problem Solving & Reasoning Multilingual Dataset
This repository contains a specialized dataset focused on logical reasoning, problem-solving, and step-by-step cognitive workflows across multiple regional languages: Pashto (ps), Arabic (ar), Farsi (fa), Sindhi (sd), and Urdu (ur).
Dataset Overview
- Languages: Pashto (پښتو), Arabic (العربية), Farsi (فارسی), Sindhi (سنڌي), Urdu (اردو)
- Domain: Logical Reasoning, Problem Solving, Cognitive SFT
- License: MIT
Structure
The dataset is structured to facilitate Supervised Fine-Tuning (SFT) and alignment for Large Language Models (LLMs) to enhance their reasoning capabilities.
{
"instruction": "...",
"input": "...",
"output": "..."
}
Dataset Files
The data files are organized in the data/ directory by language:
data/pashto_chunk_01.jsonldata/arabic_chunk_01.jsonldata/farsi_chunk_01.jsonldata/sindhi_chunk_01.jsonldata/urdu_chunk_01.jsonl
Usage
You can easily load a specific configuration or data file using the Hugging Face datasets library:
from datasets import load_dataset
# To load the Pashto data file example:
dataset = load_dataset(
"nassimjp/problem_solving-reasoning-pashto",
data_files="data/pashto_chunk_01.jsonl"
)
print(dataset["train"][0])
License
This project is released under the MIT License.
