Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01garrethlee /comprehensive-arithmetic-problemstext1M<n<10M0 likes4.5k downloads5mo agoHugging Face02EleutherAI /arithmeticA small battery of 10 tests that involve asking language models a simple arithmetic problem in natural language.text10K<n<100K5 likes4.3k downloads4y agoHugging Face03garrethlee /comprehensive-arithmetic-problems-carriestext1M<n<10M0 likes3.7k downloads2y agoHugging Face04ec75hash /qwen36-arithmetic-readouts Arithmetic Intermediate Readout Sensitivity on Qwen3.6-27B Arithmetic intermediate detection with a Jacobian lens depends sharply on the prompt token being read and the numeral forms accepted by the scorer. Across 105 order-of-operations items, the recorded hosted-lens responses contain the intermediate at rank 1 on 48 items at the trailing space, versus 4 at the preceding token. On the 25 held-out items with two-digit intermediates, rank-1 detection falls from 13 to 3 when… See the full description on the dataset page: https://huggingface.co/datasets/ec75hash/qwen36-arithmetic-readouts.tabularothern<1K0 likes1.2k downloads25d agoHugging Face05Yujivus /nanochat-climbmix-arithmetic-base10 nanochat ClimbMix + Base-10 Arithmetic This dataset contains the first 170 shuffled ClimbMix training shards used by nanochat's speedrun. The deterministic base-10 arithmetic corpus is mixed into shards 00000..00149; the final 20 train shards are unchanged web-only padding. The original validation shard (shard_06542.parquet) is also copied unchanged. Arithmetic corpus Family Examples a + b = c (all ordered pairs 0..2000, two exposures) 8,008,002 a + b… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base10.texttext-generation10M<n<100M0 likes861 downloads2mo agoHugging Face06garrethlee /simple-arithmetic-problemstext100K<n<1M2 likes604 downloads2y agoHugging Face07mib-bench /arithmetic_additiontabular10K<n<100K0 likes569 downloads1y agoHugging Face08Yujivus /nanochat-climbmix-arithmetic-base7 nanochat ClimbMix + Arithmetic: base-7 numeral world This is a deterministic base-7 rendering of Yujivus/nanochat-climbmix-arithmetic-base10. It preserves the exact shard names, row order, document order, arithmetic-document placement, and non-numeric text of the source dataset. Transformation rule Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10 integer and rendered in base 7. Leading zeros are preserved as a prefix; signs, punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base7.texttext-generation10M<n<100M0 likes462 downloads2mo agoHugging Face09Emulated-Inc /countdown-arithmetic-training-pool Countdown arithmetic training pool Arithmetic puzzles of the Countdown kind: a handful of source numbers, a target, and the job of writing an expression over the four operations that reaches the target, using each source number at most once and not having to use them all. A set generated for this pool and three public datasets read at the pinned revisions named below, laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/countdown-arithmetic-training-pool.tabulartext-generation1M<n<10M0 likes422 downloads29d agoHugging Face10Yujivus /nanochat-climbmix-arithmetic-base6 nanochat ClimbMix + Arithmetic: base-6 numeral world This is a deterministic base-6 rendering of Yujivus/nanochat-climbmix-arithmetic-base10. It preserves the exact shard names, row order, document order, arithmetic-document placement, and non-numeric text of the source dataset. Transformation rule Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10 integer and rendered in base 6. Leading zeros are preserved as a prefix; signs, punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base6.texttext-generation10M<n<100M0 likes420 downloads2mo agoHugging Face11cerebras /TAT-QA-Arithmetic-CoT Dataset Information A Chain of Thought (CoT) version of the TAT-QA arithmetic dataset (hosted at https://huggingface.co/datasets/nvidia/ChatQA-Training-Data). The dataset was synthetically generated by prompting Llama3 70B Instruct. The dataset was created as part of our work on Cerebras DocChat - a document-based conversational Q&A model. We observed that initial iterations of our model frequently made errors on arithmetic tasks (such as ConvFinQA) because it was trained on… See the full description on the dataset page: https://huggingface.co/datasets/cerebras/TAT-QA-Arithmetic-CoT.text1K<n<10K6 likes417 downloads2y agoHugging Face12mib-bench /arithmetic_subtractiontabular10K<n<100K0 likes399 downloads1y agoHugging Face13SagheerLab /Arithmetic-Reasoning SagheerLab/Arithmetic-Reasoning A high-quality synthetic arithmetic and elementary mathematics reasoning dataset for training and evaluating small language models - not an "ultimate math" claim, but a clean, verified, tiered reasoning dataset where every answer is programmatically checked. This dataset was built to train 100M-ish models that benefit disproportionately from clean, unambiguous examples. At 50M examples (45M train / 2.5M val / 2.5M test, ~5GB parquet) it is… See the full description on the dataset page: https://huggingface.co/datasets/SagheerLab/Arithmetic-Reasoning.tabulartext-generation10M<n<100M3 likes357 downloads2mo agoHugging Face14shoumenchougou /RWKV-7-ArithmeticRWKV-7-Arithmetic-0.1B 加减法运算模型的训练和测试数据集。 该模型实现基础加减法运算和加减法方程求解功能,能够处理整数部分为 1-12 位、小数部分为 0-6 位的数值,支持中英文数字、全半角格式以及大小写字符的多种表示形式,可实现基础加减法运算和加减法方程求解功能。 训练数据集说明 以下是我们使用的加减法训练数据类型,共包含 30000587 33000147 条单轮加减法 QA 数据,约 1B(1014434168) token。 数据文件名 数据条数 数据说明 示例 ADD_4M 3997733 1. 使用‘全角’、‘中文数字’、‘大写中文数字’随机替换整个数字2. 运算符附近有 1~2 个随机空格3. 含简单自然语言描述/自然语言噪声 {"text": "User: 249476576 减 796580834 还剩多少?\n\nAssistant: -547104258"} ADD_2M 1999673 1. 使用‘全角’、‘中文数字’、‘大写中文数字’随机替换整个数字2. 运算符附近有 1~2… See the full description on the dataset page: https://huggingface.co/datasets/shoumenchougou/RWKV-7-Arithmetic.text10M<n<100M0 likes311 downloads1y agoHugging Face15arithmetic-circuit-overloading /synthetic-dataset-v2-3d-3M-300K-0.1-reverse-padzerotext10M<n<100M0 likes153 downloads6mo agoHugging Face16arithmetic-circuit-overloading /synthetic-dataset-v2-3d-5M-500K-0.1-reverse-padzerotext10M<n<100M0 likes147 downloads6mo agoHugging Face17arithmetic-circuit-overloading /synthetic-dataset-2d-500K-50K-0.2-padzerotext10M<n<100M0 likes146 downloads8mo agoHugging Face18arithmetic-circuit-overloading /synthetic-dataset-v2-3d-2M-200K-0.1-reverse-padzerotext10M<n<100M0 likes143 downloads6mo agoHugging Face19arithmetic-circuit-overloading /synthetic-dataset-1d-1M-100K-0.1-reversetext10M<n<100M0 likes142 downloads8mo agoHugging Face20flexitok /mod-arithmetic Modular Arithmetic Dataset Synthetic dataset of modular-arithmetic problems of the form a mod b, paired with the result and a hypothesis about the most suitable tokenizer. Tokenizer hypothesis For a mod b where b = 2^k × 5^j (no other prime factors), only the rightmost max(k, j) digits of a determine the answer, because 10^max(k,j) ≡ 0 (mod b). A tokenizer that groups digits right-to-left in chunks of that size exposes the relevant information as a single token. For all… See the full description on the dataset page: https://huggingface.co/datasets/flexitok/mod-arithmetic.tabularquestion-answering1M<n<10M0 likes133 downloads7mo agoHugging Face21arithmetic-circuit-overloading /synthetic-dataset-v2-3d-4M-400K-0.1-reverse-padzerotext10M<n<100M0 likes129 downloads6mo agoHugging Face22Lots-of-LoRAs /task087_new_operator_addsub_arithmetic Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task087_new_operator_addsub_arithmetic Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task087_new_operator_addsub_arithmetic.texttext-generation1K<n<10K0 likes126 downloads2y agoHugging Face23arithmetic-circuit-overloading /synthetic-dataset-2d-500K-50K-0.1-reverse-padzerotext10M<n<100M0 likes120 downloads8mo agoHugging Face24arithmetic-circuit-overloading /synthetic-dataset-2d-1M-100K-0.1-padzerotext10M<n<100M0 likes119 downloads8mo agoHugging Face25ardauzunoglu /wsum-arithmetictext100K<n<1M0 likes117 downloads7mo agoHugging Face26donoway /deepmind-math-arithmetictext10M<n<100M1 likes116 downloads11mo agoHugging Face27arithmetic-circuit-overloading /synthetic-dataset-1d-1M-100K-0.1-padzerotext10M<n<100M0 likes116 downloads8mo agoHugging Face28ESITime /tram-arithmetic-responsestext10K<n<100K0 likes115 downloads1y agoHugging Face29narendarcodes /adaption-sec-financial-arithmetic-dataset SEC Financial Arithmetic Dataset — Adaption AutoScientist Challenge Powered by Adaptive Data — Adaption Labs What This Dataset Teaches This dataset trains a model to extract numbers from SEC filing tables and execute verified multi-step arithmetic — every answer is cross-checked against a gold reasoning program: Task Source Example Table Variable Extraction FinQA "From this 10-K table, extract 2021 and 2022 revenue values" Multi-Step Arithmetic… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/adaption-sec-financial-arithmetic-dataset.textquestion-answering1K<n<10K1 likes113 downloads3mo agoHugging Face30thoughtworks /arithmetic-sorl-data Arithmetic SoRL Data Training and evaluation data for the SoRL Arithmetic Interpretability Study. Small transformers trained on integer addition/subtraction, with SoRL to externalize carry/borrow circuits as explicit abstraction tokens. Reference: Quirke et al., "Understanding Addition and Subtraction in Transformers" (2024). Paper: arXiv:2402.02619 — see Table 8 for complexity classification and Section 3 for sub-task definitions. Dataset Structure Subfolder… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/arithmetic-sorl-data.tabulartext-generation1M<n<10M0 likes110 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.