LLM_Reasoning
Qwen3.5-40B-Claude-4.5-Opus-High-Reasoning-Thinking-uncensored-heretic-GGUFllm-jp-4-8b-thinking_imabari_qa_v4_reasoning_effort_v3llm-jp-4-8b-thinking_imabari_qa_v4_reasoning_effort_v4llm-jp-4-8b-thinking_imabari_qa_v4_reasoning_effort_v2llm-jp-4-8b-thinking_imabari_qa_v4_reasoning_effort_v1Phi-4-reasoning-plus-w4a16-llmcompressorLLMBG-Llama-3.1-8B-BG-Reasoning-v0.1-imatrix-GGUFBell-LLM-20B-Reasoning-GGUF
details_alexredna__Tukan-1.1B-Chat-reasoning-sft-COLA
Dataset Card for Evaluation run of alexredna/Tukan-1.1B-Chat-reasoning-sft-COLA
Dataset automatically created during the evaluation run of model alexredna/Tukan-1.1B-Chat-reasoning-sft-COLA on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_alexredna__Tukan-1.1B-Chat-reasoning-sft-COLA.High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning
PyReason-7k: Advanced Python Chain-of-Thought Dataset
Dataset Description
This dataset contains 7,000+ high-quality Python programming examples designed for LLM fine-tuning.
Each entry includes a detailed thought_process (Chain-of-Thought) to teach models logical reasoning before coding.
Key Features:
Chain-of-Thought: Step-by-step reasoning traces.
Error Handling: Solutions include try-except blocks and logging.
Diverse Tasks: Algorithms, API handling, Data Structures.… See the full description on the dataset page: https://huggingface.co/datasets/xTayyub/High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning.llm-fol-reasoning-eval
LLM FOL Reasoning Eval
This dataset is derived from ProverQA, a First-Order Logic reasoning benchmark designed to test the ability of large language models (LLMs) to perform structured logical reasoning.It restructures and normalizes the ProverQA development and training data into a unified, clean format suitable for evaluating chain-of-thought (CoT) and symbolic reasoning capabilities in LLMs.
Source
Original dataset: ProverQA: A First-Order Logic Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/MinaGabriel/llm-fol-reasoning-eval.twi-llm-reasoning-dataset-1k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Twi Reasoning Dataset
A Twi (Akan) translation of the Multilingual-Thinking… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-llm-reasoning-dataset-1k.details_alexredna__TinyLlama-1.1B-Chat-v1.0-reasoning-v2-dpo
Dataset Card for Evaluation run of alexredna/TinyLlama-1.1B-Chat-v1.0-reasoning-v2-dpo
Dataset automatically created during the evaluation run of model alexredna/TinyLlama-1.1B-Chat-v1.0-reasoning-v2-dpo on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_alexredna__TinyLlama-1.1B-Chat-v1.0-reasoning-v2-dpo.NLP-Course-LLM-Reasoning-Eval-May2025
Overview of LLM Reasoning Eval Dataset
This dataset contains evaluation of multiple large language models (LLMs) over 918 MCQ reasoning questions created by 184 students.
Each question was used to test 3 LLMs (each 3 times): GPT-4o, Claude Sonnet 3.x (3.5 or 3.7), and Deepseek R1.
The questions target various reasoning areas (i.e., Math, Logic, Temporal, Commonsense) and are included only if 3 seperate attempts (in a new session) by ChatGPT (GPT-4o) fail at giving the correct… See the full description on the dataset page: https://huggingface.co/datasets/nlpllmeval/NLP-Course-LLM-Reasoning-Eval-May2025.
