datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rocketleague-analysis
Rocket League Analysis
Local Rocket League replay analysis using Ballchasing API exports and plain DuckDB.
The report is meant to answer one practical question: what should I work on next from my saved replay sample?
Quick Start
mise install
mise run setup
mise run test
mise exec -- python scripts/analyze_scenarios.py \
--replay-dir /path/to/Rocket\ League/TAGame/Demos \
--limit 10
Start with CONTRIBUTING.md before changing the pipeline.
Replay files and… See the full description on the dataset page: https://huggingface.co/datasets/edmundmiller/rocketleague-analysis.retina-age-analysis
Retina Age Analysis Dataset
Dataset Description
This dataset contains 9,857 retinal fundus images from 5,393 patients for age prediction tasks.
Dataset Summary
Task: Age prediction from retinal fundus images
Images: 9,857 high-quality retinal images
Patients: 5,393 unique patients
Age Range: 5-97 years
Image Format: JPEG
Average Image Size: ~1 MB
Supported Tasks
Regression: Predict continuous age (5-97 years)
Classification: Predict age group (5… See the full description on the dataset page: https://huggingface.co/datasets/ramankamran/retina-age-analysis.spotify-huge-track-analysis-dataset
Spotify Track Analysis Dataset
General Description
This dataset provides a large-scale, research-oriented analytical representation of Spotify music data.
It is centered on tracks as musical recordings (track_id), while preserving explicit artist attribution as defined by Spotify’s native credit model.
Each row corresponds to a track–artist association, identified by:
a Spotify track identifier (track_id)
a credited artist name (artist_name)
A single track may appear on… See the full description on the dataset page: https://huggingface.co/datasets/GildasLeDrogoff/spotify-huge-track-analysis-dataset.source-analysis
NuBerea Source Analysis
Source-critical analysis of the Hebrew Bible, Septuagint, New Testament, Vulgate, and Second Temple literature. The dataset carries machine-generated source and tradition annotations at the verse level — the classical concerns of source criticism (documentary strata in the Old Testament, corpus structure in the New Testament, the pathway of Old Testament traditions into New Testament citation) expressed as structured data — together with semantic-domain… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/source-analysis.browsecomp-plus-selected-tools-analysis-v1
BrowseComp-Plus: Selected Tools Analysis
Side-by-side view of selected tool calls from a reference trajectory alongside the new agent trajectory conditioned on those steps.
Retrieval model: Qwen3-Embedding-8BAgent model: gpt-oss-120bRun: traj_summary_ext_selected_tools_gpt-oss-120b_seed0
Columns
Column
Description
query_id
Query identifier
rationale
GPT rationale for why these k steps were selected from the reference trajectory
selected_indices
Step indices… See the full description on the dataset page: https://huggingface.co/datasets/timchen0618/browsecomp-plus-selected-tools-analysis-v1.Instacart-Market-Basket-Analysisstudent-burnout-analysis2026
🔥 Predicting Academic Burnout: A Multivariate Analysis of Student Stressors
Exploring how financial pressure, family expectations, and social support shape burnout in university students.
Project Overview & Data Walkthrough
📋 Abstract
Academic burnout is an increasingly recognized phenomenon with far-reaching consequences for student wellbeing and performance. This study investigates the relationship between external environmental stressors —… See the full description on the dataset page: https://huggingface.co/datasets/eliel2003/student-burnout-analysis2026.septuagint-analysis
NuBerea Septuagint Textual Analysis
Curated datasets for study of the Septuagint (the ancient Greek translation of the
Hebrew Bible), part of the NuBerea corpus estate of biblical and patristic texts.
It gathers Septuagint verse texts, apparatus notes, and edition-comparison material
into a set of ready-to-load configurations.
Attribution
This dataset derives from the following upstream sources, which require attribution:
Source
License
Rahlfs… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/septuagint-analysis.multiclass-sentiment-analysis-dataset
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Sp1786/multiclass-sentiment-analysis-dataset.lxx-analysis
NuBerea Research: LXX Translation-Technique Noise Model
Quantitative study of Septuagint translation technique: verse-by-verse measurements of
where the ancient Greek translation (LXX) diverges from the Hebrew Masoretic Text, with
book-level statistical summaries. The material lets researchers distinguish a translator's
habitual working style — free versus literal rendering — from genuine textual anomalies
worth close scholarly attention, putting on a measurable footing what LXX… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/lxx-analysis.translation-analysis
NuBerea Translation Verse Texts
Verse-level texts of historical Bible translations (Clementine Vulgate, Luther Bible 1545, Matthew's Bible 1537). Part of the NuBerea curated corpus estate of biblical and historical texts.
Attribution
Upstream Data Sources
Source
License
Clementine Vulgate, NOCR
Public Domain
Luther Bible 1545, NOCR
Public Domain
Matthew's Bible 1537, Textus Receptus Bibles
Public Domain
NuBerea project. Licensed… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/translation-analysis.pseudepigrapha-analysis
NuBerea Pseudepigrapha Analysis
Derived linguistic datasets over pseudepigraphal literature, part of the NuBerea curated corpus estate. Covers the Greek and Latin witnesses of these texts along with a multilingual view across the available witness languages.
License
CC BY 4.0.
Attribution
Source
Link
License
NuBerea project
https://huggingface.co/NuBerea
CC BY 4.0
risk-analysisticker_analysis_articlesticker_analysis_pricesswebench-verified-deepseek-v4-flash-failure-analysis
SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent
Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model
driven by mini-swe-agent, graded with the official
SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the
submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative
root-cause diagnosis.
Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.student-depression-analysis
Assignment #1: EDA & Dataset
Predicting and Preventing Student Depression
Student: Amit GoodmanProgram: Economics & Entrepreneurship, Reichman University (RUNI)Date: March 2026
Project Overview
In this project, I explore the "Student Depression Dataset" to build a narrative around student well-being. By analyzing academic pressure, financial stress, and lifestyle habits, I aim to identify predictable risk factors and uncover actionable protective measures.… See the full description on the dataset page: https://huggingface.co/datasets/ag00dman/student-depression-analysis.amazon-reviews-sentiment-analysis
Dataset Card for amazon reviews for sentiment analysis
Dataset Summary
One of the most important problems in e-commerce is the correct calculation of the points given to after-sales products. The solution to this problem is to provide greater customer satisfaction for the e-commerce site, product prominence for sellers, and a seamless shopping experience for buyers. Another problem is the correct ordering of the comments given to the products. The prominence of misleading… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/amazon-reviews-sentiment-analysis.startup-Investments-analysis
📊 StartUp Investments EDA
1. Background & Objectives
This project explores a comprehensive dataset of startup investments (sourced from Crunchbase) to uncover the primary factors that predict a startup's survival and trajectory in a competitive market.
Through this Exploratory Data Analysis (EDA), we analyze historical funding data, investment rounds, and market categories to determine which variables drive specific company outcomes - namely, whether a business… See the full description on the dataset page: https://huggingface.co/datasets/lia-prop13/startup-Investments-analysis.Youtube_news_comments_analysis_ver2
YouTube 뉴스 댓글 감정 분석 — 최종 통합본
기존 분석과 2026-10-07 추가 분석을 댓글 생성일 기준으로 통합했습니다.
총 74,136,841개 댓글, 2014-04-21~2026-09-29, 관측 날짜 3,467일.
구조와 날짜
final/raw/json/YYYY/MM/01-10/news_comments.parquet
final/raw/json/YYYY/MM/11-20/news_comments.parquet
final/raw/json/YYYY/MM/21-말일/news_comments.parquet
경로와 파일 내 날짜 정렬 기준은 **dataset_date(댓글 작성일)**입니다. 윤년과 실제 말일을 반영합니다.
news_date와 기존 호환용 date는 영상 게시일입니다. 댓글 시계열 집계에 사용하지 마세요.
신규 원본 comment_created_at, video_uploaded_at은 UTC 시각입니다.… See the full description on the dataset page: https://huggingface.co/datasets/MindCastSogang/Youtube_news_comments_analysis_ver2.gretel-financial-risk-analysis-v1
gretelai/gretel-financial-risk-analysis-v1
This dataset contains synthetic financial risk analysis text generated by fine-tuning Phi-3-mini-128k-instruct on 14,306 SEC filings (10-K, 10-Q, and 8-K) from 2023-2024, utilizing differential privacy. It is designed for training models to extract key risk factors and generate structured summaries from financial documents while demonstrating the application of differential privacy to safeguard sensitive information.
This dataset showcases… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-financial-risk-analysis-v1.yeast_comparative_analysis
Yeast Comparative Analysis
Analyses that relate samples from different datasets in the BrentLab yeast
collection. Each row relates two or more samples, which are referenced by composite
identifiers of the form repo_id;config_name;sample_id, rather than describing the
samples of a single dataset. Every config here has dataset_type: comparative.
The repository currently holds one analysis, dto, and can hold others as additional
configs.
dto: dual threshold optimization… See the full description on the dataset page: https://huggingface.co/datasets/BrentLab/yeast_comparative_analysis.diabetes_eda_analysis
Diabetes Dataset — Exploratory Data Analysis (EDA)
This repository contains a diabetes-related tabular dataset and a complete Exploratory Data Analysis (EDA).The main objective of this project was to learn how to conduct a structured EDA, apply best practices, and extract meaningful insights from real-world health data.
The analysis includes correlations, distributions, group comparisons, class balance exploration, and statistical interpretations that illustrate how different… See the full description on the dataset page: https://huggingface.co/datasets/guyshilo12/diabetes_eda_analysis.digikala-sentiment-analysisXBRL_analysis
XBRL Extraction Dataset
The is the official dataset introduced in the paper FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets
Deforest-Analysis
NRT Forest-Loss Test Set for Student Analysis
This package contains the fixed held-out test split used for a study of
near-real-time forest-loss detection from four HLS observations. It is an
analysis release: it includes inputs, labels, model outputs, and visual
renders, but no checkpoints or GPU-dependent code.
The intended analyses are prediction-shape comparison, per-connected-component
performance, and seasonal performance. Do not use this test set to select model… See the full description on the dataset page: https://huggingface.co/datasets/mqraitem/Deforest-Analysis.details_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2_v2
Dataset Card for Evaluation run of deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2
Dataset automatically created during the evaluation run of model deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2.
The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2_v2.crane-analysis-cachestudent-depression-analysis
Assignment #1: EDA & Dataset
Predicting and Preventing Student Depression
Student: Amit GoodmanProgram: Economics & Entrepreneurship, Reichman University (RUNI)Date: March 2026
Project Overview
In this project, I explore the "Student Depression Dataset" to build a narrative around student well-being. By analyzing academic pressure, financial stress, and lifestyle habits, I aim to identify predictable risk factors and uncover actionable protective… See the full description on the dataset page: https://huggingface.co/datasets/urmah/student-depression-analysis.details_D2IL-Arabic-Qwen2.5-72B-Instruct-v0.1
