Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01multimodal-reasoning-lab /Zebra-CoT Zebra‑CoT A diverse large-scale dataset for interleaved vision‑language reasoning traces. Dataset Description Zebra‑CoT is a diverse large‑scale dataset with 182,384 samples containing logically coherent interleaved text‑image reasoning traces across four major categories: scientific reasoning, 2D visual reasoning, 3D visual reasoning, and visual logic & strategic games. Dataset Structure Each example in Zebra‑CoT consists of: Problem statement:… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-reasoning-lab/Zebra-CoT.imageany-to-any100K<n<1M81 likes12k downloads8mo agoHugging Face02WildEval /ZebraLogicPaper: https://huggingface.co/papers/2502.01100 Arxiv: https://arxiv.org/abs/2502.01100 Citation @article{zebralogic2025, title={ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning}, author={Bill Yuchen Lin and Ronan Le Bras and Kyle Richardson and Ashish Sabharwal and Radha Poovendran and Peter Clark and Yejin Choi}, year={2025}, url={https://arxiv.org/abs/2502.01100}, } @article{dziri2024faith, title={Faith and fate: Limits of transformers on… See the full description on the dataset page: https://huggingface.co/datasets/WildEval/ZebraLogic.text1K<n<10K18 likes5.6k downloads2y agoHugging Face03SightLinks /YOLO-OBB-Zebra-Crossings-Datasetimage1K<n<10K0 likes1.2k downloads2y agoHugging Face04RuoliuYang /textlatent_zebra_thinkmorph_armAB Text-Latent (Arm A) vs All-Latent (Arm B) — Zebra-CoT + ThinkMorph 35638 samples/arm, 18 categories. Schema = ULVR/williamium style (sample_id, category, source_dataset, question, answer, input_image, intermediate_image_N, num_intermediate_steps, messages_json). armA_text_latent: real decoded text CoT + latent visual blocks (intermediate_image_1..3). armB_render_latent: reasoning text RENDERED to images, all-latent baseline (intermediate_image_1..17). messages_json = full Monet… See the full description on the dataset page: https://huggingface.co/datasets/RuoliuYang/textlatent_zebra_thinkmorph_armAB.image100K<n<1M0 likes939 downloads3mo agoHugging Face05allenai /ZebraLogicBenchtext1K<n<10K27 likes804 downloads2y agoHugging Face06alexandrainst /multi-zebra-logic Dataset Card for the MultiZebraLogic dataset This dataset includes zebra puzzles in 39 European and 5 non-European languages and in two sizes: 2x3 and 4x5. It can be used for evaluating logical reasoning ability. The data has been generated using the code in this repo. Dataset Details Dataset Description Zebra puzzles are a type of constraint satisfaction problem. They describe a number of objects, N_objects, that each have attributes… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/multi-zebra-logic.texttext-generation100K<n<1M1 likes692 downloads2mo agoHugging Face07yrlyrl /spatial-mmcot-zebra_multihop Spatial MMCoT v1 · zebra_multihop Zebra-CoT multi-hop object counting over rendered 3D scenes (primitive objects on textured ground under a sky). Many traces open with a viewpoint change, so the first target is often a novel view. The final thought re-examines the last target image, but on some rows it only says it counts the objects and never states the number; the count is then only in <answer>. Released rows carry 2 to 5 target images. 2,057 traces with more target images… See the full description on the dataset page: https://huggingface.co/datasets/yrlyrl/spatial-mmcot-zebra_multihop.imagevisual-question-answering1K<n<10K0 likes498 downloads7d agoHugging Face08yrlyrl /spatial-mmcot-zebra_tetris Spatial MMCoT v1 · zebra_tetris Zebra-CoT Tetris: polyomino puzzles of three kinds: apply a sequence of transformations to a shape, tile a shape with a set of pieces, and fill the grid outside a shape. The options are drawn in the input image, so the answer is a letter. Each step has its own target image: transformation puzzles start with a redraw of the start shape, then one image per transformation; tiling puzzles start with the isolated shape, fill-the-complement puzzles with… See the full description on the dataset page: https://huggingface.co/datasets/yrlyrl/spatial-mmcot-zebra_tetris.imagevisual-question-answering10K<n<100K0 likes452 downloads9d agoHugging Face09allenai /ZebraLogicBench-privategatedtext1K<n<10K15 likes343 downloads2y agoHugging Face10yrlyrl /spatial-mmcot-zebra_jigsaw Spatial MMCoT v1 · zebra_jigsaw Zebra-CoT visual jigsaw, type-2 (single-image) rows only, on ImageNet images (mostly photographs; some are web graphics such as banners and flyer templates). Each row has one input image: the source image with its missing piece(s) greyed out, above a panel of four candidate piece sets labelled A-D. The options exist only in that image, so the answer is the letter. The target is the complete source image. Type-1 rows are excluded: three quarters of… See the full description on the dataset page: https://huggingface.co/datasets/yrlyrl/spatial-mmcot-zebra_jigsaw.imagevisual-question-answering10K<n<100K0 likes343 downloads7d agoHugging Face11chestnutlzj /Zebra-CoT-unify-styleimage10K<n<100K0 likes307 downloads10mo agoHugging Face12mistral-hackaton-2026 /zebra-cot-mistral-small-3.2-24b-preprocessed Zebra-CoT Preprocessed — Mistral Hackathon 2026 Preprocessed version of the Zebra-CoT dataset for fine-tuning Mistral-Small-3.2-24B-Instruct. Format text: formatted as [INST] question [/INST] <think> reasoning </think> answer image: PIL JPEG image for the corresponding visual task Usage Fine-tuning Mistral-Small-3.2-24B on chain-of-thought visual reasoning. Hackathon Created for Mistral Hackaton 2026 — Fine-tuning track with W&B. imagevisual-question-answering100K<n<1M0 likes243 downloads7mo agoHugging Face13dhruveshpatel /zebra-puzzlesSynthetic data for the paper [2505.05755] Insertion Language Models: Sequence Generation with Arbitrary-Position Insertions. Project page: https://dhruveshp.com/projects/ilm texttext-generation1M<n<10M0 likes200 downloads1y agoHugging Face14carbonteq /rg-zebra_puzzles-instruct-100k RLVR generated dataset Procedural rows from reasoning-gym, formatted for verl GRPO. Build metadata { "config": "/home/owais/Projects/rlvr/rlvr/configs/datasets/zebra_puzzles-instruct.yaml", "template_type": "qwen-instruct", "developer_prompt": null, "data_source": "reasoning_gym", "default_extract": "answer_tag", "train_rows": 100000, "test_rows": 4096, "train_seed": 42, "test_seed": 43, "tasks": { "zebra_puzzles": { "weight": 1… See the full description on the dataset page: https://huggingface.co/datasets/carbonteq/rg-zebra_puzzles-instruct-100k.text100K<n<1M0 likes124 downloads6mo agoHugging Face15AvinashAmballa /zebra-puzzles-sortedtext1M<n<10M0 likes113 downloads2y agoHugging Face16sapienzanlp /zebra-kb-explanations ZEBRA: Zero-Shot Example-Based Retrieval Augmentation for Commonsense Question Answering                     A retrieval augmentation framework for zero-shot commonsense question answering with LLMs. 🛠️ Installation Installation from PyPi pip install zebra-qa Installation from source git clone https://github.com/sapienzanlp/zebra.git cd zebra conda create -n zebra python==3.10 conda activate zebra pip install -e . 🚀 Quick Start… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/zebra-kb-explanations.text100K<n<1M3 likes106 downloads2y agoHugging Face17ZebraArena /ZebraArena ZebraArena Dataset accompanying the paper ZebraArena: A Diagnostic Simulation Environment for Studying Reasoning–Action Coupling in Tool-Augmented LLMs. ZebraArena is a procedurally generated diagnostic environment for studying reasoning–action coupling in tool-augmented LLMs, with controllable difficulty and a knowledge-minimal design. Each task is a partially observed Zebra (logic-grid) puzzle: a Constraint Satisfaction Problem with a unique ground-truth solution, where a subset… See the full description on the dataset page: https://huggingface.co/datasets/ZebraArena/ZebraArena.tabularquestion-answering1K<n<10K0 likes99 downloads5mo agoHugging Face18sunyiyou /math_logic_zebralogic_traintext1K<n<10K0 likes81 downloads1y agoHugging Face19nyu-dice-lab /lm-eval-results-mlabonne-Zebrafish-7B-private Dataset Card for Evaluation run of mlabonne/Zebrafish-7B Dataset automatically created during the evaluation run of model mlabonne/Zebrafish-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-mlabonne-Zebrafish-7B-private.tabular100K<n<1M0 likes72 downloads2y agoHugging Face20TTTXXX01 /logic__zebra_puzzle_dataset_200textn<1K0 likes69 downloads1y agoHugging Face21TTTXXX01 /Puzzle_Zebra_Alltext10K<n<100K0 likes69 downloads1y agoHugging Face22TTTXXX01 /Puzzle_Zebra_20Ktext10K<n<100K0 likes67 downloads1y agoHugging Face23TTTXXX01 /rlvr_logic__zebra_puzzle_1.3ktext1K<n<10K0 likes66 downloads1y agoHugging Face24WanjiaZhao /ZebraArena ZebraArena ZebraArena dataset. Configs missing1 … missing7 correspond to the number of missing clues; splits small / medium / large correspond to puzzle size. Usage from datasets import load_dataset ds = load_dataset("WanjiaAlicia/ZebraArena", "missing3", split="large") tabular1K<n<10K0 likes51 downloads5mo agoHugging Face25saurabh5 /Puzzle_Zebra_20K_completionstext1K<n<10K0 likes48 downloads1y agoHugging Face26tamewild /zebra_100 Zebra 100 Overview This synthetic micro-dataset contains 177 pure logic deduction traces (5x5 zebra puzzles), consisting of 100 training examples and 77 validation examples. It features strictly logic grid puzzles and includes zero mathematical data. The 100 training examples were randomly sampled without replacement from the train split of tamewild/instruct5, while the 77 validation examples are kept exactly the same across both datasets. It was created to… See the full description on the dataset page: https://huggingface.co/datasets/tamewild/zebra_100.textn<1K0 likes47 downloads1mo agoHugging Face27sunyiyou /math_logic_puzzles_zebralogic_level_4text1K<n<10K0 likes42 downloads1y agoHugging Face28TTTXXX01 /Puzzle_Zebratextn<1K0 likes42 downloads1y agoHugging Face29sunyiyou /math_logic_puzzles_zebralogic_level_5text1K<n<10K0 likes41 downloads1y agoHugging Face30sunyiyou /math_logic_puzzles_zebralogic_level_1textn<1K0 likes40 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.