Team Ai
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Emulated-Inc /countdown-arithmetic-training-pool Countdown arithmetic training pool Arithmetic puzzles of the Countdown kind: a handful of source numbers, a target, and the job of writing an expression over the four operations that reaches the target, using each source number at most once and not having to use them all. A set generated for this pool and three public datasets read at the pinned revisions named below, laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/countdown-arithmetic-training-pool.tabulartext-generation1M<n<10M0 likes422 downloads29d agoHugging Face02SagheerLab /Arithmetic-Reasoning SagheerLab/Arithmetic-Reasoning A high-quality synthetic arithmetic and elementary mathematics reasoning dataset for training and evaluating small language models - not an "ultimate math" claim, but a clean, verified, tiered reasoning dataset where every answer is programmatically checked. This dataset was built to train 100M-ish models that benefit disproportionately from clean, unambiguous examples. At 50M examples (45M train / 2.5M val / 2.5M test, ~5GB parquet) it is… See the full description on the dataset page: https://huggingface.co/datasets/SagheerLab/Arithmetic-Reasoning.tabulartext-generation10M<n<100M3 likes357 downloads2mo agoHugging Face03thoughtworks /arithmetic-sorl-data Arithmetic SoRL Data Training and evaluation data for the SoRL Arithmetic Interpretability Study. Small transformers trained on integer addition/subtraction, with SoRL to externalize carry/borrow circuits as explicit abstraction tokens. Reference: Quirke et al., "Understanding Addition and Subtraction in Transformers" (2024). Paper: arXiv:2402.02619 — see Table 8 for complexity classification and Section 3 for sub-task definitions. Dataset Structure Subfolder… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/arithmetic-sorl-data.tabulartext-generation1M<n<10M0 likes110 downloads6mo agoHugging Face04foxycuter /column-arithmetic-ru-synthetic Column Arithmetic RU Dataset Синтетический датасет для обучения модели сложению и вычитанию в столбик. Splits train.jsonl: основное обучение eval.jsonl: holdout-оценка hard.jsonl: трудные случаи с длинными переносами и займами Hard cases included 9999+1 10000+9999 9090+1010 55555+55555 10999+2 1234+8766 1000-7 10000-9999 50005-49999 8000-1 10101-909 100000-1 99009+991 12000-3456 700000+300001 1002003-998877 Current release status… See the full description on the dataset page: https://huggingface.co/datasets/foxycuter/column-arithmetic-ru-synthetic.tabulartext-generation1K<n<10K1 likes98 downloads5mo agoHugging Face05ChrisMcCormick /basic-arithmetic Basic Arithmetic Difficulty-balanced arithmetic dataset (addition, subtraction, multiplication, division) for evaluating and fine-tuning language models. Problems are classified into four difficulty tiers (easy, medium_easy, medium_hard, hard) based on Qwen2.5-0.5B-Instruct performance. Includes 10k training samples, 200 validation, and 400 test (in-domain + out-of-domain phrasings). Splits config split rows what default train 10,000 training set… See the full description on the dataset page: https://huggingface.co/datasets/ChrisMcCormick/basic-arithmetic.tabulartext-generation10K<n<100K0 likes84 downloads2mo agoHugging Face06sujitpandey /k-10-maths-arithmetic-60k k-10-maths-arithmetic-60k Synthetic, original expository text aligned to the K-10 (CBSE/NCERT-style) curriculum. Rows: 59,986 Total words: 36,319,351 Subject(s): Mathematics Rows per grade: 1: 8,319, 10: 2,730, 2: 8,318, 3: 8,320, 4: 8,319, 5: 8,312, 6: 7,513, 7: 2,730, 8: 2,730, 9: 2,695 Fields Field Type Description text string The generated passage subject string Subject name grade int Grade level word_count int Number of words in text tabulartext-generation10K<n<100K0 likes57 downloads3d agoHugging Face07vmal /3-digit-arithmetic-scratchpad-traces Contents Split Rows train 100,000 validation 4,000 test 4,000 total 108,000 Splits are prompt-disjoint — no expression appears in more than one split, and commutative swaps and trace keys are de-duplicated across splits to prevent split leakage. Operation Rows × 32,000 ÷ 32,000 + 22,000 − 22,000 Operands lie in [−999, 999]. Division answers use a fixed DDD.ddd form (round-half-up to three decimals); division by zero is an atomic <nan>.… See the full description on the dataset page: https://huggingface.co/datasets/vmal/3-digit-arithmetic-scratchpad-traces.tabulartext-generation100K<n<1M0 likes24 downloads3mo agoHugging Face08Amartya77 /Arithmetic Recursive Arithmetic Transformer training frames There was no static training file. The model sampled integers online every step. This dataset replays that sampler with the same Python RNGs used in train_recursive: RNG Seed What it draws mix_rng seed + 91 task mix, operand lengths, integers offset_rng seed + 17 Position Coupling origin seed = 42 for every run. Each train() call resets both RNGs, so later finetunes are not a continuation of earlier streams.… See the full description on the dataset page: https://huggingface.co/datasets/Amartya77/Arithmetic.tabulartext-generation1M<n<10M0 likes17 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.