Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yasumorishima /mlb-stats MLB Shared Stats Season-level MLB tables, refreshed weekly by yasumorishima/mlb-data-pipeline and published here as Parquet. One file per table at the repository root. Read the freshness column before using a table. Not everything here is current, and the rows in a stale table look exactly like the rows in a fresh one. Tables File Source Seasons Refreshed sc_batter_exitvelo.parquet Baseball Savant 2015– weekly sc_pitcher_exitvelo.parquet Baseball… See the full description on the dataset page: https://huggingface.co/datasets/yasumorishima/mlb-stats.tabular100K<n<1M2 likes1.8k downloads3d agoHugging Face02MLBricks /cosmopedia MLBricks maintained mirror \n> Upstream: HuggingFaceTB/cosmopedia @ 0ae6ec63f91742bd2d1eaef4f02232c55d719385 \n> Upstream license: Apache-2.0 \n> MLBricks maintains this repository for stable Studio presets and does not claim ownership of the source dataset.\n\n--- dataset_info: config_name: auto_math_text features: name: prompt dtype: string name: text_token_length dtype: int64 name: text dtype: string name: seed_data dtype: string name: format dtype: string name: audience dtype: string… See the full description on the dataset page: https://huggingface.co/datasets/MLBricks/cosmopedia.text10M<n<100M0 likes823 downloads1mo agoHugging Face03SmartStake /mlb-player-props SmartStake MLB Player Prop Odds and Results (2026) Minute-by-minute MLB player prop odds from ~75 sportsbooks and exchanges over the 2026 season, with the graded outcome of each prop attached. Every row is one book's price for one selection at one minute. This is the raw material behind the study "Sharpest Sportsbooks for MLB Player Props". Coverage Odds: late March 2026 through early July 2026. Graded outcomes: March through June (games that had settled at… See the full description on the dataset page: https://huggingface.co/datasets/SmartStake/mlb-player-props.tabulartime-series-forecasting100M<n<1B4 likes499 downloads3mo agoHugging Face04Hugnd-UIT /ML-Based-Malicious-Package-Detection DySec A Machine Learning-Based Dynamic Analysis For Detecting Malicious Packages In PyPI Ecosystem Overview Malicious python packages make software supply chains vulnerable by exploiting trust in open-source repositories like PyPI.Lack of real-time behavioral monitoring renders metadata inspection and static code analysis inadequate against advanced attack strategies such as typosquatting, covert remote access activation, and dynamic payload generation. To… See the full description on the dataset page: https://huggingface.co/datasets/Hugnd-UIT/ML-Based-Malicious-Package-Detection.text0 likes440 downloads7d agoHugging Face05GEM /mlb_data_to_textThe MLB dataset for data to text generation contains Major League Baseball games statistics and their human-written summaries.texttable-to-text10K<n<100K4 likes319 downloads4y agoHugging Face06treychase /mlb-daily-reporttextn<1K0 likes249 downloads7h agoHugging Face07jab13 /mlb-statcast-datasettabular1M<n<10M1 likes227 downloads11mo agoHugging Face08elcax1 /mlb-statcasttabular1M<n<10M0 likes172 downloads3h agoHugging Face09michaelmallari /mlb-statcast-batterstabular1K<n<10K0 likes151 downloads3y agoHugging Face10mlburnham /Pol_NLI Dataset Card for "Pol_NLI" More Information needed Citation To cite the paper introducing this dataset, please use: @misc{burnham2024politicaldebateefficientzeroshot, title={Political DEBATE: Efficient Zero-shot and Few-shot Classifiers for Political Text}, author={Michael Burnham and Kayla Kahn and Ryan Yank Wang and Rachel X. Peng}, year={2024}, eprint={2409.02078}, archivePrefix={arXiv}, primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/mlburnham/Pol_NLI.texttext-classification100K<n<1M6 likes146 downloads2y agoHugging Face11Coyevans /mlb-polymarket-kalshi-matched-book-sample MLB Cross-Venue Matched Book — Free Sample One full MLB game (Arizona Diamondbacks @ Minnesota Twins, 2026-06-21), with Polymarket and Kalshi prices aligned tick-for-tick and the settled outcome labeled on every row. This is a single-game sample of a larger archive. The point it proves: across the whole game, both venues priced the Twins' win probability within ~1¢ of each other on average — climbing together from ~0.10 to ~0.99 as Minnesota (the eventual winner) pulled away.… See the full description on the dataset page: https://huggingface.co/datasets/Coyevans/mlb-polymarket-kalshi-matched-book-sample.tabular1K<n<10K1 likes139 downloads3mo agoHugging Face12super-dainiu /ml-bench ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code 📖 Paper • 🚀 Github Page • 🦙 GitHub ML-Bench is a novel dual-setup benchmark designed to evaluate Large Language Models (LLMs) and AI agents in generating repository-level code for machine learning tasks. The benchmark consists of 9,641 examples from 169 diverse tasks across 18 GitHub machine learning repositories. This dataset contains the following fields:… See the full description on the dataset page: https://huggingface.co/datasets/super-dainiu/ml-bench.tabular10K<n<100K2 likes137 downloads2y agoHugging Face13MLBtrio /genz-slang-dataset Dataset Details This dataset contains a rich collection of popular slang terms and acronyms used primarily by Generation Z. It includes detailed descriptions of each term, its context of use, and practical examples that demonstrate how the slang is used in real-life conversations. The dataset is designed to capture the unique and evolving language patterns of GenZ, reflecting their communication style in digital spaces such as social media, text messaging, and online forums. Each… See the full description on the dataset page: https://huggingface.co/datasets/MLBtrio/genz-slang-dataset.texttext-generation1K<n<10K52 likes134 downloads2y agoHugging Face14Syntrex /2026_MLB_Modeltabular10M<n<100M1 likes122 downloads6mo agoHugging Face15Neonlightzz /mlb-player-props SmartStake MLB Player Prop Odds and Results (2026) Minute-by-minute MLB player prop odds from ~75 sportsbooks and exchanges over the 2026 season, with the graded outcome of each prop attached. Every row is one book's price for one selection at one minute. This is the raw material behind the study "Sharpest Sportsbooks for MLB Player Props". Coverage Odds: late March 2026 through early July 2026. Graded outcomes: March through June (games that had settled at… See the full description on the dataset page: https://huggingface.co/datasets/Neonlightzz/mlb-player-props.tabulartime-series-forecasting100M<n<1B1 likes117 downloads3mo agoHugging Face16MLBricks /openwebmath-1b MLBricks OpenWebMath 1B MLBricks curated datasetUpstream: open-web-math/open-web-math @ fde8ef8de2300f5e778f56261843dab89f230815Source config: defaultSource split: trainLicense: ODC-By 1.0 MLBricks maintains this dataset as a stable Studio preset and does not claim ownership of the underlying source material. Edition Destination: MLBricks/openwebmath-1b Rows: 469,065 GPT-2 tokens: 1,000,000,000 Main Studio column: text Source-rights notice… See the full description on the dataset page: https://huggingface.co/datasets/MLBricks/openwebmath-1b.texttext-generation100K<n<1M0 likes98 downloads1mo agoHugging Face17finnnnnnnnnnnn /mlb-play-by-plays-v1text100K<n<1M0 likes96 downloads1y agoHugging Face18mlburnham /PoliStance_AffectDataset for training an entailment classifier to recognize approval/disapproval of politicians. Documents are Tweets from Kawintiranon (2022), the MTSD dataset, as well as Tweets and sentences taken weekly newsletters for select politicians from the 115th, 116th, and 117th congress. Documents are triple coded -- once from the original compilers of the dataset, once from GPT-4, and a third time to adjudicate discrepancies between the two. Twitter handles from politicians in the dataset have… See the full description on the dataset page: https://huggingface.co/datasets/mlburnham/PoliStance_Affect.tabularzero-shot-classification10K<n<100K0 likes80 downloads2y agoHugging Face19MLBricks /tinystories MLBricks maintained mirror \n> Upstream: roneneldan/TinyStories @ f54c09fd23315a6f9c86f9dc80f725de7d8f9c64 \n> Upstream license: CDLA-Sharing-1.0 \n> MLBricks maintains this repository for stable Studio presets and does not claim ownership of the source dataset.\n\n--- license: cdla-sharing-1.0 task_categories: text-generation language: en Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in the following paper:… See the full description on the dataset page: https://huggingface.co/datasets/MLBricks/tinystories.text1M<n<10M0 likes79 downloads1mo agoHugging Face20TEAMREBOOTT-AI /SciCap-MLBCAP MLBCAP: Multi-LLM Collaborative Caption Generation in Scientific Documents 📄 PaperMLBCAP has been accepted for presentation at AI4Research @ AAAI 2025. 🎉 📌 Introduction Scientific figure captioning is a challenging task that demands contextually accurate descriptions of visual content. Existing approaches often oversimplify the task by treating it as either an image-to-text conversion or text summarization problem, leading to suboptimal results. Furthermore, commonly… See the full description on the dataset page: https://huggingface.co/datasets/TEAMREBOOTT-AI/SciCap-MLBCAP.imagetext-generation10K<n<100K19 likes68 downloads2y agoHugging Face21PengruiFu /ml-bench ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code 📖 Paper • 🚀 Github Page • 🦙 GitHub ML-Bench is a novel dual-setup benchmark designed to evaluate Large Language Models (LLMs) and AI agents in generating repository-level code for machine learning tasks. The benchmark consists of 9,641 examples from 169 diverse tasks across 18 GitHub machine learning repositories. This dataset contains the following fields:… See the full description on the dataset page: https://huggingface.co/datasets/PengruiFu/ml-bench.tabular10K<n<100K0 likes66 downloads8mo agoHugging Face22SmartStake /mlb-lineup-reactions SmartStake MLB Lineup Reactions (2026) Full-resolution sportsbook odds movements around every Underdog MLB lineup post of the 2026 season, for the six player-prop markets a batting-order change moves. Every row is one book's price for one selection at one moment, in a window around a lineup drop. This is the raw material behind the study "MLB Lineup Drops", a companion to SmartStake MLB Player Prop Odds and Results. Coverage Events: 2,456 Underdog MLB lineup… See the full description on the dataset page: https://huggingface.co/datasets/SmartStake/mlb-lineup-reactions.tabulartime-series-forecasting10M<n<100M1 likes64 downloads3mo agoHugging Face23MLBenyamin /items_prompts_litetext10K<n<100K0 likes59 downloads2mo agoHugging Face24mlburnham /global_warming_stance_entailment Dataset Card for "global_warming_stance_entailment" More Information needed tabular1K<n<10K0 likes48 downloads2y agoHugging Face25mlburnham /political_or_not Dataset Card for "political_or_not" More Information needed tabular1K<n<10K1 likes44 downloads2y agoHugging Face26Oronto /baseball-stats-cleaned_oddsportal_mlbtabular10K<n<100K0 likes42 downloads2y agoHugging Face27mlburnham /supreme_court_summary_entailment Dataset Card for "supreme_court_summary_entailment" More Information needed tabular10K<n<100K1 likes38 downloads2y agoHugging Face28MLBricks /fineweb-edu-1b MLBricks FineWeb-Edu 1B MLBricks curated datasetUpstream: HuggingFaceFW/fineweb-edu @ 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9Source config: sample-10BTSource split: trainLicense: ODC-By 1.0 MLBricks maintains this dataset as a stable Studio preset and does not claim ownership of the underlying source material. Edition Destination: MLBricks/fineweb-edu-1b Rows: 969,429 GPT-2 tokens: 1,000,000,000 Main Studio column: text Source-rights notice… See the full description on the dataset page: https://huggingface.co/datasets/MLBricks/fineweb-edu-1b.texttext-generation100K<n<1M0 likes35 downloads1mo agoHugging Face29michaelmallari /mlb-statcast-pitcherstabularn<1K1 likes34 downloads3y agoHugging Face30mlburnham /PoliStance_Affect_QTThis dataset contains quote tweets that have been hand labeled for stance towards a politician. Quote tweets are a particularly challenging classification task because they contain multiple (often contradictory) expressions from multiple authors. Twitter handles from politicians in the dataset have been replaced by their name. Be aware that "rt @realdonaldtrump You're a liar!" has been replaced with "rt trump You're a liar!", And means the author is retweeting Trump calling someone else a liar… See the full description on the dataset page: https://huggingface.co/datasets/mlburnham/PoliStance_Affect_QT.tabular1K<n<10K0 likes34 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.