datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
financial-excel-modeling-sftc4-en-html-with-training_metadata_allmodeling_valuation_knowledge
Finance Training Data Repository
A curated collection of financial modeling courses, materials, and resources designed to serve as training data for building a finance industry knowledge base.
Repository Structure
Finance_Training_Data/
├── 01_Financial_Statement_Modeling/ # 3-statement modeling fundamentals
├── 02_DCF_Modeling/ # Discounted cash flow valuation
├── 03_Trading_Comps/ # Comparable company analysis
├──… See the full description on the dataset page: https://huggingface.co/datasets/financeindustryknowledgeskills/modeling_valuation_knowledge.mmlu-winogrande-afr
Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments
Authors:
Tuka Alhanai tuka@ghamut.com, Adam Kasumovic adam.kasumovic@ghamut.com, Mohammad Ghassemi ghassemi@ghamut.com, Aven Zitzelberger aven.zitzelberger@ghamut.com, Jessica Lundin jessica.lundin@gatesfoundation.org, Guillaume Chabot-Couture Guillaume.Chabot-Couture@gatesfoundation.org
This HuggingFace Dataset contains the human-translated… See the full description on the dataset page: https://huggingface.co/datasets/Institute-Disease-Modeling/mmlu-winogrande-afr.financial-statement-modeling-sft-dpo-2026
📈 Enterprise Financial AI, SEC 10-K & Valuation Modeling SFT/DPO Dataset (2026)
High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step arithmetic Chain-of-Thought (<thought>) reasoning chains for fine-tuning LLMs (Llama-3.3, Qwen-2.5-Coder, DeepSeek-R1-Distill, Mistral) into Wall Street Equity Research Associates, M&A Valuation Modelers, and Senior Forensic Auditors.
📊 Dataset Architecture & Highlights
Multi-Turn… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/financial-statement-modeling-sft-dpo-2026.red_teaming_reward_modeling_pairwise_no_as_an_ai
Dataset Card for "red_teaming_reward_modeling_pairwise_no_as_an_ai"
More Information needed
design-bench
SciModelingBench Design-Bench Data
Canonical, provenance-tracked observations for scientific modeling and design Tasks.
GitHub
·
Python Package
·
Documentation
·
Organization
This repository stores the scientific observation layer used by the
SciModelingBench Design-Bench suite. The Python package supplies validators,
Agent-visible Protocols, trusted Objectives, submission contracts, and Task
metrics. Data and evaluation logic… See the full description on the dataset page: https://huggingface.co/datasets/sci-modeling-bench/design-bench.red_teaming_reward_modeling_pairwise
Dataset Card for "red_teaming_reward_modeling_pairwise"
More Information needed
Towards-Joint-Modeling-of-Dialogue-Response-and-Speech-Synthesis-based-on-Large-Language-Modelsharegpt_reward_modeling_pairwise_no_as_an_ai
Dataset Card for "sharegpt_reward_modeling_pairwise_no_as_an_ai"
More Information needed
Diffusion-Reward-Modeling-for-Text-Rendering-Dataset
🖼️ Text-to-Image Rendering Dataset
A dataset of 14k text prompts for image generation with text rendering evaluation
📚 Dataset Overview
This dataset contains 14,000 text prompts specifically designed for:
Image generation with text rendering
Evaluating text preservation in generated images
Training diffusion models for better text rendering
Each prompt comes with:
Pre-extracted target text for rendering
5 Stable Diffusion 3 generated latents (70k total)
Dual… See the full description on the dataset page: https://huggingface.co/datasets/leffff/Diffusion-Reward-Modeling-for-Text-Rendering-Dataset.modeling_datac4-en-html-with-metadatareward_modeling_dataset
Dataset Card for "reward_modeling_dataset"
More Information needed
dolly_reward_modeling_pairwise
Dataset Card for "dolly_reward_modeling_pairwise"
More Information needed
sharegpt_reward_modeling_pairwise
Dataset Card for "sharegpt_reward_modeling_pairwise"
More Information needed
gpteacher_reward_modeling_pairwise
Dataset Card for "gpteacher_reward_modeling_pairwise"
More Information needed
repro-impact-influence-modeling-for-open-set-time-series-anomaly-detection-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-rethinking-genomic-modeling-through-optical-character-recognition-agent-traces
OpticalDNA reproduction — Codex agent trace
This dataset contains the raw Codex JSONL session trace for the ICML 2026
reproduction of Rethinking Genomic Modeling Through Optical Character
Recognition.
Published Trackio logbook
Paper page
Challenge instructions
Agent Trace Viewer announcement
The JSONL is uploaded directly from the matching ~/.codex/sessions entry, as
recommended by the Agent Trace Viewer. It captures the reproduction work,
Hugging Face Jobs audit, poster… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/repro-rethinking-genomic-modeling-through-optical-character-recognition-agent-traces.website_metadata_c4The dataset is in the form of a json lines file with 1,20,000 examples, where an example consists of text (extracted from C4 English dataset) and metadata fields (website description extracted from Wikipedia).
Example:
{
"text": "US10289222B2 - Handling of touch events in a browser environment - Google Patents\nHandling of touch events in a browser environment Download PDF\nUS10289222B2\nUS10289222B2 US13/857,848 US201313857848A US10289222B2 US 10289222 B2 US10289222 B2 US 10289222B2 US… See the full description on the dataset page: https://huggingface.co/datasets/bs-modeling-metadata/website_metadata_c4.behavior-modeling-benchmarkmodeling
Dataset Card for "modeling"
More Information needed
churn_modeling_datasetChurnn_modeling_datasetilm_deprocess-modeling-dataset
Process Modeling Dataset
This dataset contains pairs of natural language process descriptions and their corresponding POWL (Process-Oriented Workflow Language) code implementations. It is designed for fine-tuning language models to translate informal process descriptions into formal process models.
Dataset Structure
The dataset consists of two splits:
train: Training examples for model fine-tuning
validation: Validation examples for monitoring training progress
Each… See the full description on the dataset page: https://huggingface.co/datasets/maghwa/process-modeling-dataset.ilm_esmasked_language_modeling_for_Telugu_languageT0_modeling_experimentation-canary-errors-v1
T0_modeling_experimentation-canary-errors-v1
Per-target, per-horizon prediction errors for the T0 selection placebo (canary, 10 seeds x 2 sigma_z x 5 arms x 2 targets x 14 methods x 26 horizons).
Dataset Info
Rows: 72800
Columns: 23
Columns
Column
Type
Description
sigma_z
Value('float64')
loading score std σ_z of the panel's DGP (1.0 or 2.0)
seed
Value('int64')
panel seed; the same seed is the same panel (y, m, q, u) in every arm… See the full description on the dataset page: https://huggingface.co/datasets/meganrichards3/T0_modeling_experimentation-canary-errors-v1.T0_modeling_experimentation-canary-config-v1
T0_modeling_experimentation-canary-config-v1
Everything needed to reproduce the run: full config, commit, runner and score source, and the diff against main.
Dataset Info
Rows: 1
Columns: 7
Columns
Column
Type
Description
run
Value('large_string')
local output directory
stage
Value('large_string')
canary
commit
Value('large_string')
git commit (+dirty = uncommitted changes, see git_diff_main)
config_json
Value('large_string')
the… See the full description on the dataset page: https://huggingface.co/datasets/meganrichards3/T0_modeling_experimentation-canary-config-v1.
