Team Ai
20 results

LLM_Reasoning

open-llm-leaderboard-old /details_alexredna__Tukan-1.1B-Chat-reasoning-sft-COLA Dataset Card for Evaluation run of alexredna/Tukan-1.1B-Chat-reasoning-sft-COLA Dataset automatically created during the evaluation run of model alexredna/Tukan-1.1B-Chat-reasoning-sft-COLA on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_alexredna__Tukan-1.1B-Chat-reasoning-sft-COLA.6 likes415 downloads3y agoHugging FacexTayyub /High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning PyReason-7k: Advanced Python Chain-of-Thought Dataset Dataset Description This dataset contains 7,000+ high-quality Python programming examples designed for LLM fine-tuning. Each entry includes a detailed thought_process (Chain-of-Thought) to teach models logical reasoning before coding. Key Features: Chain-of-Thought: Step-by-step reasoning traces. Error Handling: Solutions include try-except blocks and logging. Diverse Tasks: Algorithms, API handling, Data Structures.… See the full description on the dataset page: https://huggingface.co/datasets/xTayyub/High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning.text-generation1K<n<10K15 likes310 downloads10mo agoHugging FaceMinaGabriel /llm-fol-reasoning-eval LLM FOL Reasoning Eval This dataset is derived from ProverQA, a First-Order Logic reasoning benchmark designed to test the ability of large language models (LLMs) to perform structured logical reasoning.It restructures and normalizes the ProverQA development and training data into a unified, clean format suitable for evaluating chain-of-thought (CoT) and symbolic reasoning capabilities in LLMs. Source Original dataset: ProverQA: A First-Order Logic Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/MinaGabriel/llm-fol-reasoning-eval.tabulartext-classification1K<n<10K8 likes276 downloads1y agoHugging Faceghanaopenai /twi-llm-reasoning-dataset-1k This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github Twi Reasoning Dataset A Twi (Akan) translation of the Multilingual-Thinking… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-llm-reasoning-dataset-1k.texttext-generationn<1K9 likes186 downloads4mo agoHugging Faceopen-llm-leaderboard-old /details_alexredna__TinyLlama-1.1B-Chat-v1.0-reasoning-v2-dpo Dataset Card for Evaluation run of alexredna/TinyLlama-1.1B-Chat-v1.0-reasoning-v2-dpo Dataset automatically created during the evaluation run of model alexredna/TinyLlama-1.1B-Chat-v1.0-reasoning-v2-dpo on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_alexredna__TinyLlama-1.1B-Chat-v1.0-reasoning-v2-dpo.0 likes167 downloads3y agoHugging Facenlpllmeval /NLP-Course-LLM-Reasoning-Eval-May2025 Overview of LLM Reasoning Eval Dataset This dataset contains evaluation of multiple large language models (LLMs) over 918 MCQ reasoning questions created by 184 students. Each question was used to test 3 LLMs (each 3 times): GPT-4o, Claude Sonnet 3.x (3.5 or 3.7), and Deepseek R1. The questions target various reasoning areas (i.e., Math, Logic, Temporal, Commonsense) and are included only if 3 seperate attempts (in a new session) by ChatGPT (GPT-4o) fail at giving the correct… See the full description on the dataset page: https://huggingface.co/datasets/nlpllmeval/NLP-Course-LLM-Reasoning-Eval-May2025.textn<1K13 likes137 downloads1y agoHugging Face