Team Ai
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01leanpolish-anon /lean-proof-compression LeanPolish: Verified Supervision for Lean Proof Compression A dataset of Lean 4 proof rewrite pairs produced by LeanPolish, a kernel-verified proof-shortening tool. Every accepted (original, replacement) pair was kernel-checked under Lean 4.21.0 with Mathlib v4.21.0 before emission, and the rewritten file was re-elaborated end-to-end by a separate out-of-process verifier. The dataset is suitable for training models that learn to compress, simplify, or select proof tactics, and… See the full description on the dataset page: https://huggingface.co/datasets/leanpolish-anon/lean-proof-compression.tabulartext-generation10K<n<100K1 likes554 downloads14d agoHugging Face02kfdong /STP_Lean_0320This is an updated version of the final training dataset of Self-play Theorem Prover as described in the paper STP: Self-play LLM Theorem Provers with Iterative Conjecturing and Proving. This dataset includes: Extracted examples from mathlib4, Generated correct proofs of statements in LeanWorkbook, Generated correct proofs of conjectures proposed by our model during self-play training. tabulartext-generation1M<n<10M4 likes203 downloads2y agoHugging Face03kfdong /STP_LeanThis is the final training dataset of Self-play Theorem Prover as described in the paper STP: Self-play LLM Theorem Provers with Iterative Conjecturing and Proving. This dataset includes: Extracted examples from mathlib4, Generated correct proofs of statements in LeanWorkbook, Generated correct proofs of conjectures proposed by our model during self-play training. tabulartext-generation1M<n<10M0 likes113 downloads2y agoHugging Face04AlignmentResearch /math-lean-hackable-rollouts Math Lean Hackable Rollouts This dataset contains 2,241 labeled multi-turn rollouts from a GRPO run on deliberately hackable Lean 4 theorem-proving tasks. The policy was nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16. The run's weakened grader accepts proofs containing sorry; the separate oracle restores Lean's sorry check. hack_detected is true exactly when the weakened grader paid the rollout but the restored oracle rejected it. Rows without a gradeable final answer were excluded… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/math-lean-hackable-rollouts.tabulartext-generation1K<n<10K0 likes59 downloads2mo agoHugging Face05seancollins /lean-quantfinance Lean 4 Formalized Quantitative Finance & Game Theory A domain-specific Lean 4 / Mathlib corpus centered on finance and market mechanisms: 2,074 theorem records + 887 definitions, extracted from a formalization pipeline and packaged for theorem-proving research (statement, proof, tactics, premises, kernel-axiom status). This is a mechanization of largely standard applied mathematics, not new finance theory. Its value is breadth in under-formalized areas — market microstructure… See the full description on the dataset page: https://huggingface.co/datasets/seancollins/lean-quantfinance.tabulartext-generation1K<n<10K0 likes31 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.