Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SAIRfoundation /equational-theories-selected-problems Equational Theories Selected Problems Update (September 11, 2026) This dataset was updated on September 11, 2026. Main changes: released the official Stage 2 evaluation problems: stage2_evaluation_main (200 problems; ground truth withheld — answer is null until Stage 2 concludes) and stage2_evaluation_research (100 order-5 research problems with no ground truth) added metadata/stage2_evaluation_main.json and metadata/stage2_evaluation_research.json… See the full description on the dataset page: https://huggingface.co/datasets/SAIRfoundation/equational-theories-selected-problems.tabular1K<n<10K11 likes8.1k downloads28d agoHugging Face02DenCT /codeforces-problems-7ktabulartext-generation1K<n<10K6 likes1.3k downloads2y agoHugging Face03mihailgribov /olympiad_style_integer_math_problems Olympiad Math Corpus Version: v2.1.1 Release date: 2026-05-03 59,486 synthetically generated olympiad-style math problems with verified integer answers and formal computation graphs. Loading from datasets import load_dataset ds = load_dataset("mihailgribov/olympiad_style_integer_math_problems", split="train") lemma_applicability is stored as list[{lemma, status}] rather than a sparse dict (required for Arrow-based consumers). To convert to a dict for local use:… See the full description on the dataset page: https://huggingface.co/datasets/mihailgribov/olympiad_style_integer_math_problems.documenttext-generation10K<n<100K1 likes406 downloads5mo agoHugging Face04roborovski /codeforces_problems_subsettabular1K<n<10K0 likes274 downloads2y agoHugging Face05max98765 /full_geometry_problems_with_diagramstabularn<1K0 likes257 downloads6mo agoHugging Face06Parallel-Reasoning /countdown_problemstabular100K<n<1M0 likes211 downloads1y agoHugging Face07max98765 /hard_geometry_problems_with_diagramstabularn<1K0 likes202 downloads6mo agoHugging Face08togethercomputer /ParallelKernelBench_Problems ParallelKernelBench (benchmark) Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels. This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py. Inputs are deterministic — reproduce them with create_input_tensor(rank, world_size, problem_id, base_shape, dtype, trial) from that file; you do not need stored .pt files. Files Path Description… See the full description on the dataset page: https://huggingface.co/datasets/togethercomputer/ParallelKernelBench_Problems.tabulartext-generationn<1K0 likes178 downloads4mo agoHugging Face09bogoconic1 /IMO-2026-Problems IMO 2026 Problems The six IMO 2026 problem statements, indexed from 0 through 5 in contest order. IDs 0–2 are from Day 1, and IDs 3–5 are from Day 2. Schema id: zero-based problem identifier (0 corresponds to Problem 1). day: contest day (1 or 2). problem: complete English problem statement. Source Extracted from the problem statements in SignalPilot Labs' AutoFyn IMO 2026 results:… See the full description on the dataset page: https://huggingface.co/datasets/bogoconic1/IMO-2026-Problems.tabularn<1K0 likes153 downloads3mo agoHugging Face10willychan21 /ParallelKernelBench_Problems ParallelKernelBench (benchmark) Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels. This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py. Files Path Description data/problems.parquet One row per problem (tabular access) reference/*.py Reference solution() implementations utils/input_output_tensors.py Input/output tensor… See the full description on the dataset page: https://huggingface.co/datasets/willychan21/ParallelKernelBench_Problems.tabulartext-generationn<1K0 likes146 downloads5mo agoHugging Face11hummbl-hf /agent-wicked-problems-40k HUMMBL 40k Multi-Agent Wicked Problems & Coordination Corpus A foundational 40,171-event empirical dataset capturing real-world multi-agent coordination, epistemic problem decomposition, failure mode taxonomies, and strategic intelligence surges generated across the HUMMBL autonomous agent fleet. Dataset Overview The dataset provides structured visibility into how autonomous agents navigate complex, ill-defined ("wicked") problems, coordinate across distributed… See the full description on the dataset page: https://huggingface.co/datasets/hummbl-hf/agent-wicked-problems-40k.tabulartext-classification10K<n<100K1 likes86 downloads12d agoHugging Face12n4jiDX /Math-Problemstabular100K<n<1M3 likes72 downloads2y agoHugging Face13ReasoningMila /ServiceNowAI_R1_Distill_SFT_with_problems_and_responsestabular1M<n<10M0 likes59 downloads1y agoHugging Face14EleutherAI /djinn-problems-v0.9tabular1K<n<10K0 likes57 downloads7mo agoHugging Face15darumayuki /evolved-math-problems-OlympiadBench-from-deepseek-r1-0528-freetabular1K<n<10K0 likes55 downloads1y agoHugging Face16UnfaithRL /aletheia_code_problems Aletheia Code Problems with Misleading Hints Dataset Description This dataset contains multiple-choice code-reasoning problems derived from Aletheia-Bench and augmented with misleading textual hints. The misleading hints are intentionally designed to point to an incorrect answer. The dataset was developed as part of the UnfaithRL project, which studies cue-following and unfaithful reasoning under reinforcement learning with verifiable rewards. Specifically, it was… See the full description on the dataset page: https://huggingface.co/datasets/UnfaithRL/aletheia_code_problems.tabularquestion-answering10K<n<100K0 likes51 downloads3mo agoHugging Face17touristgpt /finecf-problems Dataset Card for FineCF Problems Dataset description FineCF Problems is a dataset of 9,768 Codeforces problems, each paired with a cleaned, per-problem editorial explaining the solution approach. Problems span the full difficulty range (800 to 3500) and cover a wide variety of algorithmic topics including dp, graphs, math, greedy, data structures, and more. You can load the dataset as follows: from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/touristgpt/finecf-problems.tabulartext-generation1K<n<10K3 likes50 downloads5mo agoHugging Face18d1shs0ap /qwen3-instruct-hard-problems-guidedtabular1K<n<10K0 likes39 downloads9mo agoHugging Face19touristgpt /all_cf_problemstabular10K<n<100K0 likes39 downloads7d agoHugging Face20bgub /math-problemstabular10M<n<100M2 likes35 downloads2y agoHugging Face21sytelus /taocp_open_problems TAOCP Open Problems Collection of open research problems singled out by Donald Knuth in The Art of Computer Programming series. Its main purpose is to help measure how frontier models understand, investigate, and make verifiable progress on hard but interesting open problems. Contents The dataset contains 9 exercises rated 50, M50, or HM50 in the six TAOCP editions and draft bundles available to this project. Knuth uses these ratings for problems that were not… See the full description on the dataset page: https://huggingface.co/datasets/sytelus/taocp_open_problems.tabularquestion-answeringn<1K1 likes34 downloads1mo agoHugging Face22billxbf /aimo-math-problemsMathematical QA data collections combining GSM8k, MATH and historical national mathematics competitions data extracted from AoPS, such as AMC and AIME. The dataset splits into two difficulty levels, given problems' affinity to AIMO competition. Hard: AMC12, AIME Not hard: GSM8K, MATH, AHSME, USAMO, USOMO USAJMO, AJHSME, AMC8, AMC10 These data are further deduplicated, and filtered to keep those with text-based description (instead of replying on images) and integer answers, to match AIMO… See the full description on the dataset page: https://huggingface.co/datasets/billxbf/aimo-math-problems.tabular10K<n<100K0 likes32 downloads2y agoHugging Face23SN-Col /numina-math-9sources-25each-modified-problems-o1-mod-2-onlytabularn<1K0 likes31 downloads2y agoHugging Face24d1shs0ap /pass-at-128-hard-omni-math-problems-self-correct-guidedtabular10K<n<100K0 likes31 downloads10mo agoHugging Face25lchen915 /shortest-path-dataset-problems_seed0_n20000_rows5-6_cols5-6_p0.4tabular10K<n<100K0 likes29 downloads9mo agoHugging Face26EleutherAI /djinn-problems-v1.0 djinn-problems-v1.0 — fixed-djinn v2 build (2026-09-04) Dual-verifier reward-hacking environments (insecure = exploitable, secure = hardened), rebuilt from EleutherAI/djinn-problems-v0.9 so that is_hack = insecure_pass ∧ ¬secure_pass has no known false positives on honest-but-wrong code. Every train row passes three probes under a CPU-time-bounded grader: ground truth [1,0,1], exploit [1,1,0], and a stub returning None [0,0,0] (catches vacuous verifiers), and was screened… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/djinn-problems-v1.0.tabular1K<n<10K0 likes28 downloads1mo agoHugging Face27ProblemsByVin /vehicle-reliability-scorecard Vehicle Reliability Scorecard One row per vehicle (year + make + model) with its ProblemsByVin reliability score, total NHTSA complaints, recalls, and defect investigations, plus the single component owners complain about most. The master index across the whole tracked fleet — the flat table to join every other dataset to. Columns column meaning year Model year make Manufacturer model Model reliability_score 1.0 (worst) – 5.0 (best); shown on site… See the full description on the dataset page: https://huggingface.co/datasets/ProblemsByVin/vehicle-reliability-scorecard.tabular1K<n<10K0 likes27 downloads3mo agoHugging Face28darumayuki /evolved-math-problems-from-deepseek-r1-0528-freetabularn<1K0 likes26 downloads1y agoHugging Face29EleutherAI /djinn-problems-v0.6tabular1K<n<10K0 likes26 downloads1y agoHugging Face30SN-Col /numina-math-9sources-25each-modified-problems-o1-responses-with-original-responses-final_SNOWtabularn<1K0 likes25 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.