datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen36-arithmetic-readouts
Arithmetic Intermediate Readout Sensitivity on Qwen3.6-27B
Arithmetic intermediate detection with a Jacobian lens depends sharply on the prompt token being read and the numeral forms accepted by the scorer. Across 105 order-of-operations items, the recorded hosted-lens responses contain the intermediate at rank 1 on 48 items at the trailing space, versus 4 at the preceding token. On the 25 held-out items with two-digit intermediates, rank-1 detection falls from 13 to 3 when… See the full description on the dataset page: https://huggingface.co/datasets/ec75hash/qwen36-arithmetic-readouts.arithmetic_additioncountdown-arithmetic-training-pool
Countdown arithmetic training pool
Arithmetic puzzles of the Countdown kind: a handful of source numbers, a target, and the job of
writing an expression over the four operations that reaches the target, using each source number
at most once and not having to use them all. A set generated for this pool and three public
datasets read at the pinned revisions named below, laid out twice. Train on either layer or on
both.
pool.jsonl
Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/countdown-arithmetic-training-pool.arithmetic_subtractionArithmetic-Reasoning
SagheerLab/Arithmetic-Reasoning
A high-quality synthetic arithmetic and elementary mathematics reasoning dataset for training and evaluating small language models - not an "ultimate math" claim, but a clean, verified, tiered reasoning dataset where every answer is programmatically checked.
This dataset was built to train 100M-ish models that benefit disproportionately from clean, unambiguous examples. At 50M examples (45M train / 2.5M val / 2.5M test, ~5GB parquet) it is… See the full description on the dataset page: https://huggingface.co/datasets/SagheerLab/Arithmetic-Reasoning.mod-arithmetic
Modular Arithmetic Dataset
Synthetic dataset of modular-arithmetic problems of the form a mod b,
paired with the result and a hypothesis about the most suitable tokenizer.
Tokenizer hypothesis
For a mod b where b = 2^k × 5^j (no other prime factors), only the
rightmost max(k, j) digits of a determine the answer, because
10^max(k,j) ≡ 0 (mod b). A tokenizer that groups digits right-to-left
in chunks of that size exposes the relevant information as a single token.
For all… See the full description on the dataset page: https://huggingface.co/datasets/flexitok/mod-arithmetic.arithmetic-sorl-data
Arithmetic SoRL Data
Training and evaluation data for the SoRL Arithmetic Interpretability Study.
Small transformers trained on integer addition/subtraction, with
SoRL to externalize carry/borrow circuits
as explicit abstraction tokens.
Reference: Quirke et al., "Understanding Addition and Subtraction in Transformers" (2024).
Paper: arXiv:2402.02619 — see Table 8 for complexity classification and Section 3 for sub-task definitions.
Dataset Structure
Subfolder… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/arithmetic-sorl-data.column-arithmetic-ru-synthetic
Column Arithmetic RU Dataset
Синтетический датасет для обучения модели сложению и вычитанию в столбик.
Splits
train.jsonl: основное обучение
eval.jsonl: holdout-оценка
hard.jsonl: трудные случаи с длинными переносами и займами
Hard cases included
9999+1
10000+9999
9090+1010
55555+55555
10999+2
1234+8766
1000-7
10000-9999
50005-49999
8000-1
10101-909
100000-1
99009+991
12000-3456
700000+300001
1002003-998877
Current release status… See the full description on the dataset page: https://huggingface.co/datasets/foxycuter/column-arithmetic-ru-synthetic.basic-arithmetic
Basic Arithmetic
Difficulty-balanced arithmetic dataset (addition, subtraction, multiplication,
division) for evaluating and fine-tuning language models. Problems are classified
into four difficulty tiers (easy, medium_easy, medium_hard, hard) based on
Qwen2.5-0.5B-Instruct performance. Includes 10k training samples, 200
validation, and 400 test (in-domain + out-of-domain phrasings).
Splits
config
split
rows
what
default
train
10,000
training set… See the full description on the dataset page: https://huggingface.co/datasets/ChrisMcCormick/basic-arithmetic.mod3-arithmetick-10-maths-arithmetic-60k
k-10-maths-arithmetic-60k
Synthetic, original expository text aligned to the K-10 (CBSE/NCERT-style) curriculum.
Rows: 59,986
Total words: 36,319,351
Subject(s): Mathematics
Rows per grade: 1: 8,319, 10: 2,730, 2: 8,318, 3: 8,320, 4: 8,319, 5: 8,312, 6: 7,513, 7: 2,730, 8: 2,730, 9: 2,695
Fields
Field
Type
Description
text
string
The generated passage
subject
string
Subject name
grade
int
Grade level
word_count
int
Number of words in text
numerical_reasoning_arithmetic Generated dataset for testing numerical reasoningArithmetic-XL
Arithmetic-XL
12 Million Verified Arithmetic Examples for Language Model Pretraining
"Arithmetic is one of those things that humans learn in primary
school but language models somehow manage to forget halfway through
training."
Arithmetic-XL is a large-scale arithmetic pretraining dataset
containing over 12 million procedurally generated and verified
examples designed to improve numerical reasoning in language models.
The dataset was created with one objective:… See the full description on the dataset page: https://huggingface.co/datasets/GODELEV/Arithmetic-XL.calculator-arithmeticmod7-arithmeticarithmetic-3digitarithmetic_multiplicationarithmetic_subtractiontinyzero-arithmetic-3_digit3-digit-arithmetic-scratchpad-traces
Contents
Split
Rows
train
100,000
validation
4,000
test
4,000
total
108,000
Splits are prompt-disjoint — no expression appears in more than one split,
and commutative swaps and trace keys are de-duplicated across splits to prevent
split leakage.
Operation
Rows
×
32,000
÷
32,000
+
22,000
−
22,000
Operands lie in [−999, 999]. Division answers use a fixed DDD.ddd form
(round-half-up to three decimals); division by zero is an atomic <nan>.… See the full description on the dataset page: https://huggingface.co/datasets/vmal/3-digit-arithmetic-scratchpad-traces.dataset-v2-99arithmetic-llm
arithmetic-llm
MNIST 图像 → 字节级算术 训练数据集。
用于教学向的字节级算术大模型 (byte-omni-model-zh):给定一个显式数字 (digit 0-9) 与一张 MNIST 图像 (隐式数字),预测两者之和。
输入: digit_byte + MNIST_image_bytes
输出: result_bytes
示例:
digit = "1" + image(数字 "2") → 预测 "3"
Schema
列
类型
说明
id
int
序号
image_base64
str
MNIST 图像 PNG (base64)
pixels
float[]
28×28 展平像素 (0-1), 784 维
label
int
图像数字标签 (0-9)
split
str
train (59134) / test (9943)
image_hash
str
SHA-256 逐字节指纹
phash
str
16×16 感知哈希位串 (与 label… See the full description on the dataset page: https://huggingface.co/datasets/wcpsoft/arithmetic-llm.arithmetic_additiondataset-v2-999Arithmetic
Recursive Arithmetic Transformer training frames
There was no static training file. The model sampled integers online every
step. This dataset replays that sampler with the same Python RNGs used in
train_recursive:
RNG
Seed
What it draws
mix_rng
seed + 91
task mix, operand lengths, integers
offset_rng
seed + 17
Position Coupling origin
seed = 42 for every run. Each train() call resets both RNGs, so later
finetunes are not a continuation of earlier streams.… See the full description on the dataset page: https://huggingface.co/datasets/Amartya77/Arithmetic.level_01_arithmeticarithmetic-priming-datasetarithmetic-few-shotarithmeticrlvr_task086_translated_symbol_arithmetic
