Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vid-modeling /videomme0 likes39k downloads2y agoHugging Face02moritzmiller /mental-modelingimagequestion-answering10K<n<100K0 likes6.4k downloads2y agoHugging Face03piergiuliol /financial-excel-modeling-sfttext1K<n<10K2 likes4.8k downloads5mo agoHugging Face04IDEA-FinAI /Mathematical_Modeling_Speciale_Dataset_v0.1image0 likes3.8k downloads9mo agoHugging Face05bs-modeling-metadata /c4-en-html-with-training_metadata_alltabular10K<n<100K1 likes2.4k downloads4y agoHugging Face06financeindustryknowledgeskills /modeling_valuation_knowledge Finance Training Data Repository A curated collection of financial modeling courses, materials, and resources designed to serve as training data for building a finance industry knowledge base. Repository Structure Finance_Training_Data/ ├── 01_Financial_Statement_Modeling/ # 3-statement modeling fundamentals ├── 02_DCF_Modeling/ # Discounted cash flow valuation ├── 03_Trading_Comps/ # Comparable company analysis ├──… See the full description on the dataset page: https://huggingface.co/datasets/financeindustryknowledgeskills/modeling_valuation_knowledge.documentn<1K0 likes1.9k downloads4mo agoHugging Face07BVRA /plant-distribution-modeling0 likes730 downloads2mo agoHugging Face08andersonbcdefg /reward-modeling-short-tokenized Dataset Card for "reward-modeling-short-tokenized" More Information needed 100K<n<1M4 likes389 downloads3y agoHugging Face09Institute-Disease-Modeling /mmlu-winogrande-afr Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments Authors: Tuka Alhanai tuka@ghamut.com, Adam Kasumovic adam.kasumovic@ghamut.com, Mohammad Ghassemi ghassemi@ghamut.com, Aven Zitzelberger aven.zitzelberger@ghamut.com, Jessica Lundin jessica.lundin@gatesfoundation.org, Guillaume Chabot-Couture Guillaume.Chabot-Couture@gatesfoundation.org This HuggingFace Dataset contains the human-translated… See the full description on the dataset page: https://huggingface.co/datasets/Institute-Disease-Modeling/mmlu-winogrande-afr.textquestion-answering10K<n<100K1 likes344 downloads2y agoHugging Face10beatsprom /financial-statement-modeling-sft-dpo-2026 📈 Enterprise Financial AI, SEC 10-K & Valuation Modeling SFT/DPO Dataset (2026) High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step arithmetic Chain-of-Thought (<thought>) reasoning chains for fine-tuning LLMs (Llama-3.3, Qwen-2.5-Coder, DeepSeek-R1-Distill, Mistral) into Wall Street Equity Research Associates, M&A Valuation Modelers, and Senior Forensic Auditors. 📊 Dataset Architecture & Highlights Multi-Turn… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/financial-statement-modeling-sft-dpo-2026.texttext-generationn<1K0 likes320 downloads1mo agoHugging Face11andersonbcdefg /reward-modeling-long-tokenized Dataset Card for "reward-modeling-long-tokenized" More Information needed 100K<n<1M6 likes274 downloads3y agoHugging Face12andersonbcdefg /red_teaming_reward_modeling_pairwise_no_as_an_ai Dataset Card for "red_teaming_reward_modeling_pairwise_no_as_an_ai" More Information needed text10K<n<100K10 likes250 downloads3y agoHugging Face13sci-modeling-bench /design-bench SciModelingBench Design-Bench Data Canonical, provenance-tracked observations for scientific modeling and design Tasks. GitHub &nbsp;·&nbsp; Python Package &nbsp;·&nbsp; Documentation &nbsp;·&nbsp; Organization This repository stores the scientific observation layer used by the SciModelingBench Design-Bench suite. The Python package supplies validators, Agent-visible Protocols, trusted Objectives, submission contracts, and Task metrics. Data and evaluation logic… See the full description on the dataset page: https://huggingface.co/datasets/sci-modeling-bench/design-bench.tabular1M<n<10M0 likes239 downloads3mo agoHugging Face14andersonbcdefg /red_teaming_reward_modeling_pairwise Dataset Card for "red_teaming_reward_modeling_pairwise" More Information needed text10K<n<100K9 likes190 downloads3y agoHugging Face15ADS599-Capstone /modeling_datatabular10M<n<100M0 likes178 downloads6mo agoHugging Face16bs-modeling-metadata /c4-en-html-with-metadatatabular10M<n<100M14 likes177 downloads4y agoHugging Face17luckeciano /pku-llama3.1-8b-dataset-features-gt-reward-modeling3 likes160 downloads2y agoHugging Face18andersonbcdefg /sharegpt_reward_modeling_pairwise_no_as_an_ai Dataset Card for "sharegpt_reward_modeling_pairwise_no_as_an_ai" More Information needed text10K<n<100K6 likes154 downloads3y agoHugging Face19leffff /Diffusion-Reward-Modeling-for-Text-Rendering-Dataset 🖼️ Text-to-Image Rendering Dataset A dataset of 14k text prompts for image generation with text rendering evaluation 📚 Dataset Overview This dataset contains 14,000 text prompts specifically designed for: Image generation with text rendering Evaluating text preservation in generated images Training diffusion models for better text rendering Each prompt comes with: Pre-extracted target text for rendering 5 Stable Diffusion 3 generated latents (70k total) Dual… See the full description on the dataset page: https://huggingface.co/datasets/leffff/Diffusion-Reward-Modeling-for-Text-Rendering-Dataset.tabulartext-to-image10K<n<100K10 likes152 downloads1y agoHugging Face20xinyuzhou2000 /Towards-Joint-Modeling-of-Dialogue-Response-and-Speech-Synthesis-based-on-Large-Language-Modeltext10K<n<100K9 likes148 downloads3y agoHugging Face21HumanDynamics /reward_modeling_dataset Dataset Card for "reward_modeling_dataset" More Information needed text10K<n<100K4 likes145 downloads3y agoHugging Face22andersonbcdefg /gpteacher_reward_modeling_pairwise Dataset Card for "gpteacher_reward_modeling_pairwise" More Information needed text1K<n<10K4 likes111 downloads3y agoHugging Face23andersonbcdefg /sharegpt_reward_modeling_pairwise Dataset Card for "sharegpt_reward_modeling_pairwise" More Information needed text10K<n<100K4 likes110 downloads3y agoHugging Face24zifeng-ai /clinical-trial-modeling Clinical Trial Modeling complete campaign data This dataset contains development data, unlabeled verifier inputs, and held-out labels required by the trusted grader. Keep the dataset access controlled if the evaluation labels should remain hidden from campaign participants. The HMAC pseudonymization key and raw source archives are not included. Set CLINICAL_TRIAL_MODELING_DATA_DIR to the downloaded dataset root and CLINICAL_TRIAL_MODELING_PRIVATE_LABEL_DIR to its private… See the full description on the dataset page: https://huggingface.co/datasets/zifeng-ai/clinical-trial-modeling.0 likes110 downloads15d agoHugging Face25andersonbcdefg /dolly_reward_modeling_pairwise Dataset Card for "dolly_reward_modeling_pairwise" More Information needed text10K<n<100K4 likes109 downloads3y agoHugging Face26andersonbcdefg /reward-modeling-eval-tokenized Dataset Card for "reward-modeling-eval-tokenized" More Information needed 10K<n<100K7 likes104 downloads3y agoHugging Face27bs-modeling-metadata /website_metadata_c4The dataset is in the form of a json lines file with 1,20,000 examples, where an example consists of text (extracted from C4 English dataset) and metadata fields (website description extracted from Wikipedia). Example: { "text": "US10289222B2 - Handling of touch events in a browser environment - Google Patents\nHandling of touch events in a browser environment Download PDF\nUS10289222B2\nUS10289222B2 US13/857,848 US201313857848A US10289222B2 US 10289222 B2 US10289222 B2 US 10289222B2 US… See the full description on the dataset page: https://huggingface.co/datasets/bs-modeling-metadata/website_metadata_c4.text10K<n<100K4 likes82 downloads5y agoHugging Face28jomasego /repro-impact-influence-modeling-for-open-set-time-series-anomaly-detection-traces Agent traces Agent sessions published from a Trackio Logbook. text10K<n<100K0 likes80 downloads2mo agoHugging Face29PaulAlm /GenAI_Channel_Modeling_Datasets Dataset Card — Site-Specific MIMO Channel Generation via Diffusion and Flow Matching: Fidelity, Efficiency, and Downstream Utility Link to paper: https://arxiv.org/abs/2606.20098 Authors: Sina Beyraghi, Masoud Sadeghian, Firdous Bin Ismail, Angel Lozano, Paul Almasan, and Giovanni Geraci Contact: Sina Beyraghi (mohammadsina.beyraghi@telefonica.com) Abstract This paper explores the use of generative models to synthesize high-quality… See the full description on the dataset page: https://huggingface.co/datasets/PaulAlm/GenAI_Channel_Modeling_Datasets.1 likes78 downloads4mo agoHugging Face30nielsr /repro-rethinking-genomic-modeling-through-optical-character-recognition-agent-traces OpticalDNA reproduction — Codex agent trace This dataset contains the raw Codex JSONL session trace for the ICML 2026 reproduction of Rethinking Genomic Modeling Through Optical Character Recognition. Published Trackio logbook Paper page Challenge instructions Agent Trace Viewer announcement The JSONL is uploaded directly from the matching ~/.codex/sessions entry, as recommended by the Agent Trace Viewer. It captures the reproduction work, Hugging Face Jobs audit, poster… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/repro-rethinking-genomic-modeling-through-optical-character-recognition-agent-traces.tabularn<1K0 likes76 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.