datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nbm-conus-analysis
NOAA NBM CONUS Daily Analysis (Zarr)
Daily best-estimate analysis derived from NOAA NBM (National Blend of Models)
CONUS forecasts, on the native ~2.5 km Lambert conformal grid (2345 x 1597).
Built by nbm-to-zarr, dynamical.org-style.
Variables: tmean / tmax / tmin (degC), precip (mm), srad (MJ/m2/day)
Construction: best estimate for day D = lead-day 1 of that day's 00z NBM init
Coverage: rolling backfill from 2020-10-01 (AWS NBM archive floor) to present
Layout: one standalone… See the full description on the dataset page: https://huggingface.co/datasets/nakas/nbm-conus-analysis.data_analysis
Dataset Card for "livebench/data_analysis"
LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties:
LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses.
Each question has verifiable, objective ground-truth answers, allowing hard questions to be… See the full description on the dataset page: https://huggingface.co/datasets/livebench/data_analysis.Taur_CoT_Analysis_Project___gpt-4o-2024-08-06bazaarbench-analysis-v2-gpt5-judge
BazaarBench Analysis-v2 GPT-5 High-Reasoning Judgments
This release contains the complete frozen Analysis-v2 semantic-judgment run for
BazaarBench. The judge is the exact research gateway (TRAPI) deployment gpt-5_2025-08-07 with
reasoning_effort=high, a 32,768-token output cap, the strict bound-sources wire
contract, eight semantic attempts, and the user-authorized accelerated concurrency
protocol with ceiling 256.
Completion and validation
55 / 55 independent… See the full description on the dataset page: https://huggingface.co/datasets/BazaarBench/bazaarbench-analysis-v2-gpt5-judge.Lurcher_10x
Lurcher 10x Microscopy Dataset
Dataset overview
This dataset consists of 2-D microscopy images of histologically stained 3-D structures in tissue sections through the cerebellum of 21 mouse brains. Animals are grouped into wild-type controls (n = 10) and Lurcher mutant mice (n = 11). The classification task is to distinguish Lurcher mutant mice from wild-type controls.
All images were captured at low magnification (10x) and stained with Cresyl violet, a general… See the full description on the dataset page: https://huggingface.co/datasets/USF-CS-Microscopy-Image-Analysis/Lurcher_10x.NLU-Sentiment-Analysis
SEA Sentiment Analysis
SEA Sentiment Analysis evaluates a model's ability to identify the sentiment polarity of a text. It is sampled from NusaX for Indonesian, Javanese, and Sundanese, IndicSentiment for Tamil, Wisesight Sentiment for Thai, and UIT-VSFC for Vietnamese.
Supported Tasks and Leaderboards
SEA Sentiment Analysis is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Sentiment-Analysis.Openpdf-Analysis-Recognition
Openpdf-Analysis-Recognition
The Openpdf-Analysis-Recognition dataset is curated for tasks related to image-to-text recognition, particularly for scanned document images and OCR (Optical Character Recognition) use cases. It contains over 6,900 images in a structured imagefolder format suitable for training models on document parsing, PDF image understanding, and layout/text extraction tasks.
Attribute
Value
Task
Image-to-Text
Modality
Image
Format
ImageFolder… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Openpdf-Analysis-Recognition.Taur_CoT_Analysis_Project___gpt-4o-mini-2024-07-18web3-trading-analysisThis dataset contains web3-related on-chain and off-chain data, which can be used to build quantitative models.
rocketleague-analysis
Rocket League Analysis
Local Rocket League replay analysis using Ballchasing API exports and plain DuckDB.
The report is meant to answer one practical question: what should I work on next from my saved replay sample?
Quick Start
mise install
mise run setup
mise run test
mise exec -- python scripts/analyze_scenarios.py \
--replay-dir /path/to/Rocket\ League/TAGame/Demos \
--limit 10
Start with CONTRIBUTING.md before changing the pipeline.
Replay files and… See the full description on the dataset page: https://huggingface.co/datasets/edmundmiller/rocketleague-analysis.Taur_CoT_Analysis_Project___meta-llama__Meta-Llama-3.1-8B-Instructrouting_analysis-checkpoints
routing_analysis checkpoint archive
This public dataset repository stores checkpoint files from the
routing_analysis filesystem snapshot while preserving their original paths
under routing_analysis/.
The tree routing_analysis/finetuning/phase2_full/checkpoints/ is explicitly
excluded. All other regular files classified under checkpoint directories are
included, including small code and configuration files needed to keep those
checkpoint directories complete.
Files are uploaded… See the full description on the dataset page: https://huggingface.co/datasets/lylybig/routing_analysis-checkpoints.retina-age-analysis
Retina Age Analysis Dataset
Dataset Description
This dataset contains 9,857 retinal fundus images from 5,393 patients for age prediction tasks.
Dataset Summary
Task: Age prediction from retinal fundus images
Images: 9,857 high-quality retinal images
Patients: 5,393 unique patients
Age Range: 5-97 years
Image Format: JPEG
Average Image Size: ~1 MB
Supported Tasks
Regression: Predict continuous age (5-97 years)
Classification: Predict age group (5… See the full description on the dataset page: https://huggingface.co/datasets/ramankamran/retina-age-analysis.Step-analysisTaur_CoT_Analysis_Project___microsoft__Phi-3-small-8k-instructNMR-analysisrouting_analysis-marco_mini_tam_100k_500k_checkpoints
Two marco_mini_base Tam checkpoints
This dataset contains every regular file in the original
routing_analysis/finetuning/phase2_full/checkpoints/marco_mini_base/tam_Taml_100k_full
and tam_Taml_500k_full trees, at the original relative paths. The files are
individual objects, not tar archives. manifests/ records the selected source
snapshot and provenance from the local direct-upload inventory.
tam_Taml_500k_full/checkpoint-4500 is preserved as found and does not contain… See the full description on the dataset page: https://huggingface.co/datasets/divin1234/routing_analysis-marco_mini_tam_100k_500k_checkpoints.routing_analysis-marco_nano-hau_Latn-checkpoints
marco_nano_base hau_Latn files
This dataset preserves the original relative paths of every regular file under
the ten marco_nano_base directories whose basename contains the literal
hau_Latn, plus the two matching zero-byte lock files beside them.
The scope intentionally includes hidden work data, preserved backups,
failed_edquot data, and all other regular files found inside those selected
trees. No .gitignore file or checkpoint-name exclusion rule is consulted.
The source… See the full description on the dataset page: https://huggingface.co/datasets/lylybig/routing_analysis-marco_nano-hau_Latn-checkpoints.static-analysis-evalA dataset of 76 Python programs taken from real Python open source projects (top 100 on GitHub),
where each program is a file that has exactly 1 vulnerability as detected by a particular static analyzer (Semgrep), used in the paper Patched MOA: optimizing inference for diverse software development tasks.
OpenAI used the synth-vuln-fixes and fine-tuned
a new version of gpt-4o is now the SOTA on this benchmark. More details and code is available from their repo.
More details on the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/patched-codes/static-analysis-eval.gem-analysis-unimts
gem-analysis-unimts
UniMTS pretraining datasets for fitness action recognition
数据集信息
来源路径: datasets/unimts
数据大小: 6.6 GB
用途: 健身动作识别模型训练
使用方法
from huggingface_hub import snapshot_download
# 下载数据集
snapshot_download(
repo_id="yonful/gem-analysis-unimts",
repo_type="dataset",
local_dir="./datasets/unimts"
)
或使用项目中的下载脚本:
python scripts/prepare_data.py --dataset unimts
许可证
请参考原始数据源的许可证要求。
trash-in-river-2025
Street Parade 2025 Dataset
Overview
This dataset was collected by SARA, a student initiative at ETH Zurich, to enable open research on trash presence in aquatic environments. It contains images of litter in the Limmat River in Zurich the day after the Street Parade (August 9, 2025). The dataset is intended for training and evaluating trash classification models.
Dataset summary
Collection date: August 9, 2025
Location: Kornhausbrücke, Zurich… See the full description on the dataset page: https://huggingface.co/datasets/SARA-smartphone-assisted-river-analysis/trash-in-river-2025.Taur_CoT_Analysis_Project___google__gemini-1.5-flash-001spotify-huge-track-analysis-dataset
Spotify Track Analysis Dataset
General Description
This dataset provides a large-scale, research-oriented analytical representation of Spotify music data.
It is centered on tracks as musical recordings (track_id), while preserving explicit artist attribution as defined by Spotify’s native credit model.
Each row corresponds to a track–artist association, identified by:
a Spotify track identifier (track_id)
a credited artist name (artist_name)
A single track may appear on… See the full description on the dataset page: https://huggingface.co/datasets/GildasLeDrogoff/spotify-huge-track-analysis-dataset.browsecomp-plus-selected-tools-analysis-v1
BrowseComp-Plus: Selected Tools Analysis
Side-by-side view of selected tool calls from a reference trajectory alongside the new agent trajectory conditioned on those steps.
Retrieval model: Qwen3-Embedding-8BAgent model: gpt-oss-120bRun: traj_summary_ext_selected_tools_gpt-oss-120b_seed0
Columns
Column
Description
query_id
Query identifier
rationale
GPT rationale for why these k steps were selected from the reference trajectory
selected_indices
Step indices… See the full description on the dataset page: https://huggingface.co/datasets/timchen0618/browsecomp-plus-selected-tools-analysis-v1.routing_analysis-marco_nano-dan_Latn-checkpoints
marco_nano_base dan_Latn files
This dataset preserves original routing_analysis/... paths for all
regular files under the five marco_nano_base directories whose basename
contains dan_Latn. There are no separately matching regular files.
Every file inside the selected directory trees is included. No
.gitignore file or checkpoint-name exclusion rule is consulted.
The repository stores each path as a separate file. Local ownership,
permissions, and timestamps are recorded in the… See the full description on the dataset page: https://huggingface.co/datasets/lylybig8/routing_analysis-marco_nano-dan_Latn-checkpoints.lewm-vjepa21-state-moe-shared-residual-wo-tokencl-analysis
LEWM V-JEPA2.1 State-MoE Shared-Residual Analysis
This dataset contains PNG visualizations and adjacent-epoch absolute-delta
summaries for the experiment
lewm_reasoning_vjepa21_vitL_tokens_100tasks_stateMoE_sharedResidual_pair_topkExcess_flat_headaware_woTokenCL_modify.
Contents
metadata.csv: parsed stage, epoch, layer, condition, scope, and image paths.
gallery_manifest.json: manifest consumed by the companion Static HTML Space.
summary.json: aggregate image… See the full description on the dataset page: https://huggingface.co/datasets/zoeloopy/lewm-vjepa21-state-moe-shared-residual-wo-tokencl-analysis.source-analysis
NuBerea Source Analysis
Source-critical analysis of the Hebrew Bible, Septuagint, New Testament, Vulgate, and Second Temple literature. The dataset carries machine-generated source and tradition annotations at the verse level — the classical concerns of source criticism (documentary strata in the Old Testament, corpus structure in the New Testament, the pathway of Old Testament traditions into New Testament citation) expressed as structured data — together with semantic-domain… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/source-analysis.routing_analysis-marco_nano-swh_Latn-checkpoints
marco_nano_base swh_Latn files
This dataset preserves original routing_analysis/... paths for every regular
file under the three marco_nano_base directories whose basename contains
swh_Latn, plus the matching top-level hidden .lock file.
The scope includes every regular file inside the selected directory trees.
No .gitignore or checkpoint-name exclusion rule is consulted. The zero-byte
.lock file is included because the user requested every matching file.
The repository stores… See the full description on the dataset page: https://huggingface.co/datasets/lylybig8/routing_analysis-marco_nano-swh_Latn-checkpoints.qwen36-runtime-analysis-3090ti
Qwen3.6 Runtime Analysis on RTX 3090 Ti
This dataset contains a local runtime viability analysis for Qwen3.6 27B MTP and Qwen3.6 35B-A3B quantized candidates on a single RTX 3090 Ti 24GB system.
Open index.html for the formatted report with charts and separated tables for:
long stress tests: 10K input + up to 20K output
smoke/load tests: short context viability probes
Key artifacts:
index.html / analysis.html: formatted analysis report
summary.csv: flattened benchmark result… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwen36-runtime-analysis-3090ti.india-tb-missed-cases-analysis
India TB Missed Cases Analysis & Living Model (2025)
🌟 Project Overview
This repository hosts a comprehensive, multi-method analytical framework designed to estimate and understand the "missing" millions of Tuberculosis (TB) cases in India. By integrating Bayesian statistics, Dimensionality Reduction (PCA), and Causal Inference (DAG), this project provides a high-resolution view of TB detection determinants across Indian states.
Core Analytical Pillars:… See the full description on the dataset page: https://huggingface.co/datasets/hssling/india-tb-missed-cases-analysis.
