datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Krea-2-Raw_samplesThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
The images were created in ComfyUI with the
bf16 version
of… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples.hubbench
HubBench 1.4.0
One Blobfish-authored, oracle-proven benchmark family per Harbor Hub professional-domain cluster. Every task is an employee decision worked over a dependent chain of evidence — never a lookup — against mock stateful tools over an isolated SQLite world. The agent reaches the world only through its public surfaces (MCP over streamable HTTP, a terminal tool CLI, a REST API, and a web console); a deterministic verifier (HubScore) grades the finished world from… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/hubbench.Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/SBMM75/Krea-2-Raw_samples_Best_of.semikongbench-100
SemiKongBench-100
One hundred independently authored synthetic semiconductor operations tasks: ten workflows across ten frozen fabs. Each task includes 20 agent-visible assets, a task-local SQLite world, 38 provider-shaped tools, an oracle trajectory, before/after snapshots, and a deterministic 100-point verifier.
The model leaderboard is intentionally empty until a complete version-pinned 100-task model run exists. Release qualification executes oracle, replay, and six… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/semikongbench-100.Crop-Recommendation-Parameters
🌱 Crop Recommendation Dataset
A machine learning dataset for crop recommendation based on soil properties and environmental conditions. The dataset contains measurements of essential soil nutrients and climatic parameters, along with the crop label that is suitable for those conditions.
This dataset can be used for machine learning classification, agricultural analytics, decision-support systems, and smart farming applications.
📌 Dataset Overview
Property… See the full description on the dataset page: https://huggingface.co/datasets/Samarth-27/Crop-Recommendation-Parameters.Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples_Best_of.llm-cost-same-prompt
Measured per-call LLM cost — same prompt, every model
Vendors publish prices per million tokens. Nobody publishes what one call actually costs, because
that depends on how many tokens the model chooses to emit — and on the same question models differ by
more than an order of magnitude. One model finishes a JSON extraction in 23 tokens; another writes 300.
This dataset sends a fixed set of prompts to every model at temperature 0, every night, and records
the cost computed from… See the full description on the dataset page: https://huggingface.co/datasets/mario0369/llm-cost-same-prompt.ts-satfire
Dataset Card for TS-SatFire
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
The TS-SatFire dataset is a comprehensive multi-temporal remote sensing dataset designed to cover the entire life cycle of wildfires. It provides a unified framework to support three critical and interconnected wildfire monitoring tasks: active fire detection, daily burned area… See the full description on the dataset page: https://huggingface.co/datasets/SamuelWu318/ts-satfire.practice-radar-behavioral-health-npi-sample
New behavioral-health organization NPIs — weekly NPPES sample
A 15-row public sample from a weekly, reproducible selection of newly enumerated Type 2 behavioral-health organizations in the U.S. Centers for Medicare & Medicaid Services National Plan and Provider Enumeration System (NPPES).
Edition at a glance
Measured period: July 6–12, 2026
New Type 2 organizations screened: 2,722
Behavioral-health organizations selected: 486
States and territories represented:… See the full description on the dataset page: https://huggingface.co/datasets/unitedideas/practice-radar-behavioral-health-npi-sample.recalldb-product-recalls-sample
RecallDB — U.S. Product Recall Database (Sample)
Full dataset: recalldb.dataengineered.io · $49 one-time snapshot → Buy on Stripe · the same sample on Kaggle
128,936 official recalls · 294,683 recalled products · CPSC · FDA · FSIS · NHTSA · USCG · 100% source-linked
RecallDB is a normalized, provenance-tracked dataset of official U.S. federal product recalls. It joins five official source families into one relational model: CPSC consumer products, NHTSA vehicles, FDA/openFDA… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/recalldb-product-recalls-sample.point-in-time-us-equity-fundamentals-sample
Tradevo Data — honest point-in-time US equity fundamentals
Fundamentals with filed-date stamps, so a backtest only sees what was public — and restatements are flagged, not silently applied.
A deliberately small public proof pack of point-in-time US equity fundamentals, built from SEC EDGAR.
Every value is stamped with the date it first became public (first_filed), so a join that
filters by first_filed <= as_of only sees what was knowable on that date — and later
revisions are… See the full description on the dataset page: https://huggingface.co/datasets/Tradevodata/point-in-time-us-equity-fundamentals-sample.SAMTOR_Novel_Target_Designs_GA-II
SAMTOR — SAM-competitive de novo designs (Technetium GA-II)
196 small molecules across two generations, generated de novo by the Technetium TC-43.ai engine (GA-II) and conditioned on the S-adenosylmethionine (SAM) pocket of human SAMTOR.
Each molecule was constructed against this pocket rather than selected from a compound library — docking (AutoDock Vina) came afterwards, to place and score the generated molecules in the site. With no approved drug, clinical candidate or… See the full description on the dataset page: https://huggingface.co/datasets/Tc-43/SAMTOR_Novel_Target_Designs_GA-II.sam3-piglife-labels
SAM3 Pseudo-Labels for Pig Monitoring (Based on PigLife)
[!WARNING]
No Images Included (Copyright Restrictions)
Due to the copyright and licensing restrictions of the original PigLife dataset by UIUC, this repository DOES NOT contain any original images.
This dataset provides only the pseudo-labels (annotations) generated by the Segment Anything Model 3 (SAM3) in YOLO format.
To use these annotations, you must legally acquire the original images directly from… See the full description on the dataset page: https://huggingface.co/datasets/marcosvmf/sam3-piglife-labels.AniGen-Sample-Dataset
AniGen Sample Data
This directory is a compact example subset of the AniGen training dataset.
What Is Included
10 examples
10 unique raw assets
Full cross-modal files for each example
A subset metadata.csv with 10 rows
The retained directory layout follows the core structure of the reference test set:
raw/
renders/
renders_cond/
skeleton/
voxels/
features/
metadata.csv
statistics.txt
latents/ (encoded by the trained slat auto-encoder)
ss_latents/ (encoded by the… See the full description on the dataset page: https://huggingface.co/datasets/VAST-AI/AniGen-Sample-Dataset.zomato_delivery_EDA📹 Video walkthrough:
Zomato Delivery Operations — EDA & Dataset
Dataset Overview
Real-world delivery data from Zomato operations across multiple Indian cities,
covering courier attributes, weather conditions, traffic density, GPS coordinates,
and delivery outcomes.
Source
Kaggle — saurabhbadole/zomato-delivery-operations-analytics-dataset
Original size
45,584 rows × 20 columns
Final size
38,964 rows × 22 columns
Target variable… See the full description on the dataset page: https://huggingface.co/datasets/Samruddhi-Desai/zomato_delivery_EDA.whiskydb-fine-spirits-sample
🥃 WhiskyDB — Fine Spirits & Whisky Dataset (Free Sample)
Full dataset: whiskydb.dataengineered.io · $49 one-time (or $49 / month with the monthly refresh) → Buy once · Subscribe · the same sample on Kaggle
A free sample of WhiskyDB: a structured, relational dataset of whiskies and fine spirits built entirely from open, legally accessible public sources — government label registries (US TTB COLA), corporate registries (UK Companies House), the EU eAmbrosia GI register, Open… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/whiskydb-fine-spirits-sample.sample-community-dataset
Field
Type†
What it contains
challenge_id
integer
Unique numeric identifier for the coding challenge
challenge_slug
string
URL-friendly slug used in challenge links
challenge_name
string
Human-readable challenge title
challenge_body
string
Full challenge description (HTML/Markdown) including input/output, examples, etc.
challenge_kind
string
High-level content type (e.g., code, game)
challenge_preview
string
One-sentence teaser shown in listings
challenge_category
string… See the full description on the dataset page: https://huggingface.co/datasets/hackerrank/sample-community-dataset.skilled-commercial-work-egocentric-sample
🛠️ Skilled Commercial Work — Egocentric Video Dataset (Sample)
This dataset is part of a larger collection of egocentric activity datasets by Verbose Tech Labs LLP. If you want the full dataset, or want access to more categories? Get in touch with us:
📞 Phone: +91 7672 000 500
💬 WhatsApp: +91 7672 000 500
📧 Email: Hello@VerboseTechLabs.com
🌐 Website: VerboseTechLabs.com
🔗 More datasets: kaggle.com/verbosetechlabsllp
Dataset Summary
First-person… See the full description on the dataset page: https://huggingface.co/datasets/VerboseTechLabs/skilled-commercial-work-egocentric-sample.graded-player-props-sample
PropLine graded player props — free 14-day sample
Real player-prop prices from 25+ US sportsbooks, prediction markets and DFS apps,
each one graded against the official box score, with the opening and closing
line for every price. From PropLine.
Files
File
Sport
Markets
mlb_graded_props_14d_part1/2.csv.gz
MLB
pitcher strikeouts, batter hits, batter total bases
nfl_graded_props_14d.csv.gz
NFL
passing yards, rushing yards, receiving yards, receptions… See the full description on the dataset page: https://huggingface.co/datasets/propline/graded-player-props-sample.stacked-samsum-1024
stacked samsum 1024
Created with the stacked-booksum repo version v0.25. It contains:
Original Dataset: copy of the base dataset
Stacked Rows: The original dataset is processed by stacking rows based on certain criteria:
Maximum Input Length: The maximum length for input sequences is 1024 tokens in the longt5 model tokenizer.
Maximum Output Length: The maximum length for output sequences is also 1024 tokens in the longt5 model tokenizer.
Special Token: The dataset utilizes the… See the full description on the dataset page: https://huggingface.co/datasets/stacked-summaries/stacked-samsum-1024.electronics-assembly-egocentric-sample
🔌 Electronics Assembly — Egocentric Video Dataset (Sample)
This dataset is part of a larger collection of egocentric activity datasets by Verbose Tech Labs LLP. If you want the full dataset, or want access to more categories? Get in touch with us:
📞 Phone: +91 7672 000 500
💬 WhatsApp: +91 7672 000 500
📧 Email: Hello@VerboseTechLabs.com
🌐 Website: VerboseTechLabs.com
🔗 More datasets: kaggle.com/verbosetechlabsllp
Dataset Summary
First-person point-of-view… See the full description on the dataset page: https://huggingface.co/datasets/VerboseTechLabs/electronics-assembly-egocentric-sample.roasterdb-specialty-coffee-sample
☕ RoasterDB — Specialty Coffee Dataset (Free Sample)
Full dataset: roasterdb.dataengineered.io · $49 one-time → Buy on Stripe · the same sample on Kaggle
A free sample of RoasterDB: a structured dataset of specialty-coffee products scraped from the direct storefronts of curated artisan roasters worldwide, with tasting notes normalized to the Specialty Coffee Association (SCA) Flavor Wheel and a source URL on every record so any fact can be re-verified.
This sample contains 100… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/roasterdb-specialty-coffee-sample.floradb-houseplants-care-sample
🌿 FloraDB — Houseplant Care & Pet-Toxicity Dataset (Free Sample)
Full dataset: floradb.dataengineered.io · $49 one-time → Buy on Stripe · the same sample on Kaggle
A free sample of FloraDB: a structured dataset that turns subjective houseplant care advice — "bright indirect light", "water when dry" — into quantitative engineering metrics (Lux thresholds, watering-day intervals, temperature and humidity ranges), joined to ASPCA dog/cat toxicity and grounded on the GBIF… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/floradb-houseplants-care-sample.shopify-products-catalog-sample
Shopify Products Catalog with Price History (DTC Stores) — Free Sample
A free sample of a Shopify e-commerce catalog dataset built from public, login-free products.json endpoints of independent DTC (direct-to-consumer) Shopify stores.
This sample build covers 386 stores / 533,479 SKU rows (2026-08-01 snapshot). Full-run scale on record: 4,100+ stores / 6.18M SKUs, refreshed with weekly snapshots that also yield per-variant price-change and availability-change events.
➡️ Full… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/shopify-products-catalog-sample.product_hscodeJudgmentBench
JudgmentBench
JudgmentBench is an expert-annotated legal evaluation dataset for studying how different feedback protocols recover quality differences in open-ended legal work product. The dataset contains 30 real-world legal tasks, model-generated outputs at three constructed quality levels, rubric scores from practicing lawyers, pairwise comparative judgments from practicing lawyers, GPT-5.4 and GPT-5.4-mini autograder annotations for the same completed study assignments… See the full description on the dataset page: https://huggingface.co/datasets/Sami2305341176/JudgmentBench.polymarket-kalshi-scoresync-orderbook-sample
Polymarket x Kalshi Orderbook Archive - Free Score-Synced Sample
A free excerpt of the ZenHodl Polymarket & Kalshi Historical Orderbook Archive,
published so you can VERIFY the two non-reconstructable features before you buy --
the things competitors (Telonex, PolymarketData) do not ship:
Per-row score-sync - every Kalshi row carries the live game state next to the quote:
home_score, away_score, score_diff, period, time_remaining, game_state.
Cross-venue - the same games are… See the full description on the dataset page: https://huggingface.co/datasets/Coyevans/polymarket-kalshi-scoresync-orderbook-sample.crypto-5s-market-data-adausdc-sample
ADA/USDC High-Frequency Market Microstructure Data
Free 7-Day Sample
This repository provides a free 7-day sample of a much larger privately collected high-frequency cryptocurrency market dataset.
The complete historical archive contains millions of market snapshots, with data collection starting in December 2025, across 12 crypto/USDC markets.
The public ADA/USDC sample contains:
81,579 market snapshots
97 columns
7 days of continuous historical data
20 bid + 20… See the full description on the dataset page: https://huggingface.co/datasets/rfab85/crypto-5s-market-data-adausdc-sample.us-business-locations-samples
LocationLists — US business location samples
Ten real rows from each of 30 of the 726 datasets published at locationlists.com, a catalog of 43,101,581 verified US business locations: manufacturer dealer networks, retail and restaurant chains, licensed trade contractors, healthcare providers, nonprofits and membership directories. Every dataset is compiled from the brand's own store locator or the official government registry, re-checked weekly, and sold as a flat CSV with no… See the full description on the dataset page: https://huggingface.co/datasets/locationlists/us-business-locations-samples.oil037-sample
OIL-037 — Synthetic Regulatory Compliance Dataset (Sample)
A schema-identical preview of OIL-037, the XpertSystems.ai synthetic
regulatory compliance dataset for upstream, midstream, and downstream oil &
gas operations. The full product covers 8,500 facilities across 16 regulatory
frameworks (OSHA PSM, EPA CAA/CWA, API RP 754 / 1173, ISO 14001 / 45001,
PHMSA, BSEE, GHGRP, SOX, NERC CIP, IEC 62443, NIST, SEC Climate, Corporate
ESG). This sample is the generator's sample mode (500… See the full description on the dataset page: https://huggingface.co/datasets/xpertsystems/oil037-sample.
