Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01garrethlee /comprehensive-arithmetic-problemstext1M<n<10M0 likes8.5k downloads5mo agoHugging Face02garrethlee /comprehensive-arithmetic-problems-carriestext1M<n<10M0 likes6.5k downloads2y agoHugging Face03EleutherAI /arithmeticA small battery of 10 tests that involve asking language models a simple arithmetic problem in natural language.text10K<n<100K5 likes5.1k downloads4y agoHugging Face04Yujivus /nanochat-climbmix-arithmetic-base10 nanochat ClimbMix + Base-10 Arithmetic This dataset contains the first 170 shuffled ClimbMix training shards used by nanochat's speedrun. The deterministic base-10 arithmetic corpus is mixed into shards 00000..00149; the final 20 train shards are unchanged web-only padding. The original validation shard (shard_06542.parquet) is also copied unchanged. Arithmetic corpus Family Examples a + b = c (all ordered pairs 0..2000, two exposures) 8,008,002 a + b… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base10.texttext-generation10M<n<100M0 likes2.4k downloads2mo agoHugging Face05ec75hash /qwen36-arithmetic-readouts Arithmetic Intermediate Readout Sensitivity on Qwen3.6-27B Arithmetic intermediate detection with a Jacobian lens depends sharply on the prompt token being read and the numeral forms accepted by the scorer. Across 105 order-of-operations items, the recorded hosted-lens responses contain the intermediate at rank 1 on 48 items at the trailing space, versus 4 at the preceding token. On the 25 held-out items with two-digit intermediates, rank-1 detection falls from 13 to 3 when… See the full description on the dataset page: https://huggingface.co/datasets/ec75hash/qwen36-arithmetic-readouts.tabularothern<1K0 likes1.2k downloads22d agoHugging Face06Yujivus /nanochat-climbmix-arithmetic-base7 nanochat ClimbMix + Arithmetic: base-7 numeral world This is a deterministic base-7 rendering of Yujivus/nanochat-climbmix-arithmetic-base10. It preserves the exact shard names, row order, document order, arithmetic-document placement, and non-numeric text of the source dataset. Transformation rule Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10 integer and rendered in base 7. Leading zeros are preserved as a prefix; signs, punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base7.texttext-generation10M<n<100M0 likes830 downloads2mo agoHugging Face07garrethlee /simple-arithmetic-problemstext100K<n<1M2 likes783 downloads2y agoHugging Face08mib-bench /arithmetic_additiontabular10K<n<100K0 likes588 downloads1y agoHugging Face09Yujivus /nanochat-climbmix-arithmetic-base6 nanochat ClimbMix + Arithmetic: base-6 numeral world This is a deterministic base-6 rendering of Yujivus/nanochat-climbmix-arithmetic-base10. It preserves the exact shard names, row order, document order, arithmetic-document placement, and non-numeric text of the source dataset. Transformation rule Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10 integer and rendered in base 6. Leading zeros are preserved as a prefix; signs, punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base6.texttext-generation10M<n<100M0 likes472 downloads2mo agoHugging Face10mib-bench /arithmetic_subtractiontabular10K<n<100K0 likes424 downloads1y agoHugging Face11SagheerLab /Arithmetic-Reasoning SagheerLab/Arithmetic-Reasoning A high-quality synthetic arithmetic and elementary mathematics reasoning dataset for training and evaluating small language models - not an "ultimate math" claim, but a clean, verified, tiered reasoning dataset where every answer is programmatically checked. This dataset was built to train 100M-ish models that benefit disproportionately from clean, unambiguous examples. At 50M examples (45M train / 2.5M val / 2.5M test, ~5GB parquet) it is… See the full description on the dataset page: https://huggingface.co/datasets/SagheerLab/Arithmetic-Reasoning.tabulartext-generation10M<n<100M3 likes362 downloads1mo agoHugging Face12shoumenchougou /RWKV-7-ArithmeticRWKV-7-Arithmetic-0.1B 加减法运算模型的训练和测试数据集。 该模型实现基础加减法运算和加减法方程求解功能,能够处理整数部分为 1-12 位、小数部分为 0-6 位的数值,支持中英文数字、全半角格式以及大小写字符的多种表示形式,可实现基础加减法运算和加减法方程求解功能。 训练数据集说明 以下是我们使用的加减法训练数据类型,共包含 30000587 33000147 条单轮加减法 QA 数据,约 1B(1014434168) token。 数据文件名 数据条数 数据说明 示例 ADD_4M 3997733 1. 使用‘全角’、‘中文数字’、‘大写中文数字’随机替换整个数字2. 运算符附近有 1~2 个随机空格3. 含简单自然语言描述/自然语言噪声 {"text": "User: 249476576 减 796580834 还剩多少?\n\nAssistant: -547104258"} ADD_2M 1999673 1. 使用‘全角’、‘中文数字’、‘大写中文数字’随机替换整个数字2. 运算符附近有 1~2… See the full description on the dataset page: https://huggingface.co/datasets/shoumenchougou/RWKV-7-Arithmetic.text10M<n<100M0 likes327 downloads1y agoHugging Face13cerebras /TAT-QA-Arithmetic-CoT Dataset Information A Chain of Thought (CoT) version of the TAT-QA arithmetic dataset (hosted at https://huggingface.co/datasets/nvidia/ChatQA-Training-Data). The dataset was synthetically generated by prompting Llama3 70B Instruct. The dataset was created as part of our work on Cerebras DocChat - a document-based conversational Q&A model. We observed that initial iterations of our model frequently made errors on arithmetic tasks (such as ConvFinQA) because it was trained on… See the full description on the dataset page: https://huggingface.co/datasets/cerebras/TAT-QA-Arithmetic-CoT.text1K<n<10K6 likes284 downloads2y agoHugging Face14ESITime /tram-arithmetic-responsestext10K<n<100K0 likes207 downloads1y agoHugging Face15Skanth007 /arithmetic-logical-pointer-setARTHIMETIC-lOGICAL_POINTER_SET 0 likes196 downloads1mo agoHugging Face16flexitok /mod-arithmetic Modular Arithmetic Dataset Synthetic dataset of modular-arithmetic problems of the form a mod b, paired with the result and a hypothesis about the most suitable tokenizer. Tokenizer hypothesis For a mod b where b = 2^k × 5^j (no other prime factors), only the rightmost max(k, j) digits of a determine the answer, because 10^max(k,j) ≡ 0 (mod b). A tokenizer that groups digits right-to-left in chunks of that size exposes the relevant information as a single token. For all… See the full description on the dataset page: https://huggingface.co/datasets/flexitok/mod-arithmetic.tabularquestion-answering1M<n<10M0 likes152 downloads7mo agoHugging Face17arithmetic-circuit-overloading /results-v2imagen<1K0 likes151 downloads5mo agoHugging Face18arithmetic-circuit-overloading /synthetic-dataset-2d-500K-50K-0.1-reverse-padzerotext10M<n<100M0 likes147 downloads7mo agoHugging Face19arithmetic-circuit-overloading /synthetic-dataset-v2-3d-2M-200K-0.1-reverse-padzerotext10M<n<100M0 likes143 downloads6mo agoHugging Face20donoway /deepmind-math-arithmetictext10M<n<100M1 likes140 downloads11mo agoHugging Face21arithmetic-circuit-overloading /synthetic-dataset-v2-3d-5M-500K-0.1-reverse-padzerotext10M<n<100M0 likes135 downloads6mo agoHugging Face22arithmetic-circuit-overloading /synthetic-dataset-1d-1M-100K-0.1-reversetext10M<n<100M0 likes134 downloads8mo agoHugging Face23arithmetic-circuit-overloading /synthetic-dataset-2d-500K-50K-0.2-padzerotext10M<n<100M0 likes128 downloads7mo agoHugging Face24Lots-of-LoRAs /task087_new_operator_addsub_arithmetic Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task087_new_operator_addsub_arithmetic Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task087_new_operator_addsub_arithmetic.texttext-generation1K<n<10K0 likes127 downloads2y agoHugging Face25arithmetic-circuit-overloading /synthetic-dataset-v2-3d-3M-300K-0.1-reverse-padzerotext10M<n<100M0 likes127 downloads6mo agoHugging Face26arithmetic-circuit-overloading /synthetic-dataset-2d-1M-100K-0.1-padzerotext10M<n<100M0 likes126 downloads7mo agoHugging Face27arithmetic-circuit-overloading /synthetic-dataset-v2-1d-5M-500K-0.1-reverse-padzerotext10M<n<100M0 likes125 downloads6mo agoHugging Face28arithmetic-circuit-overloading /synthetic-dataset-v2-3d-4M-400K-0.1-reverse-padzerotext10M<n<100M0 likes121 downloads6mo agoHugging Face29arithmetic-circuit-overloading /synthetic-dataset-v2-3d-5M-500K-0.1-padzerotext10M<n<100M0 likes121 downloads6mo agoHugging Face30neurallambda /arithmetic_dataset Arithmetic Puzzles Dataset A collection of arithmetic puzzles with heavy use of variable assignment. Current LLMs struggle with variable indirection/multi-hop reasoning, this should be a tough test for them. Inputs are a list of strings representing variable assignments (c=a+b), and the output is the integer answer. Outputs are filtered to be between [-100, 100], and self-reference/looped dependencies are forbidden. Splits are named like: train_N 8k total examples of puzzles with N… See the full description on the dataset page: https://huggingface.co/datasets/neurallambda/arithmetic_dataset.text100K<n<1M0 likes116 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.