datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rocketleague-analysis
Rocket League Analysis
Local Rocket League replay analysis using Ballchasing API exports and plain DuckDB.
The report is meant to answer one practical question: what should I work on next from my saved replay sample?
Quick Start
mise install
mise run setup
mise run test
mise exec -- python scripts/analyze_scenarios.py \
--replay-dir /path/to/Rocket\ League/TAGame/Demos \
--limit 10
Start with CONTRIBUTING.md before changing the pipeline.
Replay files and… See the full description on the dataset page: https://huggingface.co/datasets/edmundmiller/rocketleague-analysis.retina-age-analysis
Retina Age Analysis Dataset
Dataset Description
This dataset contains 9,857 retinal fundus images from 5,393 patients for age prediction tasks.
Dataset Summary
Task: Age prediction from retinal fundus images
Images: 9,857 high-quality retinal images
Patients: 5,393 unique patients
Age Range: 5-97 years
Image Format: JPEG
Average Image Size: ~1 MB
Supported Tasks
Regression: Predict continuous age (5-97 years)
Classification: Predict age group (5… See the full description on the dataset page: https://huggingface.co/datasets/ramankamran/retina-age-analysis.web3-trading-analysisThis dataset contains web3-related on-chain and off-chain data, which can be used to build quantitative models.
Instacart-Market-Basket-Analysisturkish-sentiment-analysis-dataset
Dataset
This dataset contains positive , negative and notr sentences from several data sources given in the references. In the most sentiment models , there are only two labels; positive and negative. However , user input can be totally notr sentence. For such cases there were no data I could find. Therefore I created this dataset with 3 class. Positive and negative sentences are listed below. Notr examples are extraced from turkish wiki dump. In addition, added some random text… See the full description on the dataset page: https://huggingface.co/datasets/winvoker/turkish-sentiment-analysis-dataset.student-burnout-analysis2026
🔥 Predicting Academic Burnout: A Multivariate Analysis of Student Stressors
Exploring how financial pressure, family expectations, and social support shape burnout in university students.
Project Overview & Data Walkthrough
📋 Abstract
Academic burnout is an increasingly recognized phenomenon with far-reaching consequences for student wellbeing and performance. This study investigates the relationship between external environmental stressors —… See the full description on the dataset page: https://huggingface.co/datasets/eliel2003/student-burnout-analysis2026.multiclass-sentiment-analysis-dataset
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Sp1786/multiclass-sentiment-analysis-dataset.risk-analysisnlp_twitter_analysisstudent-depression-analysis
Assignment #1: EDA & Dataset
Predicting and Preventing Student Depression
Student: Amit GoodmanProgram: Economics & Entrepreneurship, Reichman University (RUNI)Date: March 2026
Project Overview
In this project, I explore the "Student Depression Dataset" to build a narrative around student well-being. By analyzing academic pressure, financial stress, and lifestyle habits, I aim to identify predictable risk factors and uncover actionable protective measures.… See the full description on the dataset page: https://huggingface.co/datasets/ag00dman/student-depression-analysis.amazon-reviews-sentiment-analysis
Dataset Card for amazon reviews for sentiment analysis
Dataset Summary
One of the most important problems in e-commerce is the correct calculation of the points given to after-sales products. The solution to this problem is to provide greater customer satisfaction for the e-commerce site, product prominence for sellers, and a seamless shopping experience for buyers. Another problem is the correct ordering of the comments given to the products. The prominence of misleading… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/amazon-reviews-sentiment-analysis.startup-Investments-analysis
📊 StartUp Investments EDA
1. Background & Objectives
This project explores a comprehensive dataset of startup investments (sourced from Crunchbase) to uncover the primary factors that predict a startup's survival and trajectory in a competitive market.
Through this Exploratory Data Analysis (EDA), we analyze historical funding data, investment rounds, and market categories to determine which variables drive specific company outcomes - namely, whether a business… See the full description on the dataset page: https://huggingface.co/datasets/lia-prop13/startup-Investments-analysis.snappfood-sentiment-analysisdigikala-sentiment-analysisXBRL_analysis
XBRL Extraction Dataset
The is the official dataset introduced in the paper FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets
llm-ideology-analysisThis dataset contains evaluations of political figures by a diverse set of Large Language Models (LLMs), such that the ideology of these LLMs can be characterized.
📝 Dataset Description
The dataset contains responses from 19 different Large Language Models evaluating 3,991 political figures, with responses collected in the six UN languages: Arabic, Chinese, English, French, Russian, and Spanish.
The evaluations were conducted using a two-stage prompting strategy to assess the… See the full description on the dataset page: https://huggingface.co/datasets/aida-ugent/llm-ideology-analysis.synthetic-sentiment-analysis-dataset-v1
Tanaos Sentiment Analysis Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate sentiment analysis systems — models that classify the sentiment expressed in text as one of five possible categories: very_negative, negative, neutral, positive or very_positive. It can be used to build sentiment analysis models for various applications, such as customer feedback analysis, social media… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-sentiment-analysis-dataset-v1.Deforest-Analysis
NRT Forest-Loss Test Set for Student Analysis
This package contains the fixed held-out test split used for a study of
near-real-time forest-loss detection from four HLS observations. It is an
analysis release: it includes inputs, labels, model outputs, and visual
renders, but no checkpoints or GPU-dependent code.
The intended analyses are prediction-shape comparison, per-connected-component
performance, and seasonal performance. Do not use this test set to select model… See the full description on the dataset page: https://huggingface.co/datasets/mqraitem/Deforest-Analysis.Amazon_Reviews_Binary_for_Sentiment_Analysis
Dataset Card for Dataset Name
The Amazon reviews polarity dataset is constructed by taking review score 1 and 2 as negative, and 4 and 5 as positive. Samples of score 3 is ignored. In the dataset, class 1 is the negative and class 2 is the positive. Each class has 1,800,000 training samples and 200,000 testing samples.
Dataset Details
Dataset Description
The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 3… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Amazon_Reviews_Binary_for_Sentiment_Analysis.student-depression-analysis
Assignment #1: EDA & Dataset
Predicting and Preventing Student Depression
Student: Amit GoodmanProgram: Economics & Entrepreneurship, Reichman University (RUNI)Date: March 2026
Project Overview
In this project, I explore the "Student Depression Dataset" to build a narrative around student well-being. By analyzing academic pressure, financial stress, and lifestyle habits, I aim to identify predictable risk factors and uncover actionable protective… See the full description on the dataset page: https://huggingface.co/datasets/urmah/student-depression-analysis.airlines-sentiment-analysisAmazon_Reviews_for_Sentiment_Analysis_fine_grained_5_classes
Dataset Card for Dataset Name
The Amazon reviews full score dataset is constructed by randomly taking 600,000 training samples and 130,000 testing samples for each review score from 1 to 5. In total there are 3,000,000 trainig samples and 650,000 testing samples.
Dataset Details
Dataset Description
The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 3 columns in them, corresponding to class index (1 to 5)… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Amazon_Reviews_for_Sentiment_Analysis_fine_grained_5_classes.airbnb-global-market-analysis
Global Airbnb Market Analysis: Pricing and Host Dynamics
Repository Contents
File
Description
airbnb_top_cities_final.csv
The final, cleaned dataset used for this analysis (normalized to USD).
Airbnb_Market_Analysis_EDA.ipynbThe full Python notebook containing all cleaning code and visualizations.
Google Colab Notebook
Direct link to the live interactive research environment.
README.md
This document, providing the project overview and key findings.… See the full description on the dataset page: https://huggingface.co/datasets/rotemvahava/airbnb-global-market-analysis.student-burnout-analysis-2026airline-satisfaction-analysis
Airline Passenger Satisfaction – EDA Report
This project analyzes the Airline Passenger Satisfaction Dataset, containing 103,904 rows and 25 columns describing passenger demographics, flight information, and service ratings.The goal is to understand which factors influence satisfaction, identify important service features,and compare satisfaction between different traveler types and flight classes.
Dataset Overview
The dataset includes:
Passenger demographics (age… See the full description on the dataset page: https://huggingface.co/datasets/drukeroni/airline-satisfaction-analysis.sinhala-sentiment-analysissentiment-analysis-tweetYelp_Reviews_for_Sentiment_Analysis_fine_grained_5_classes
Dataset Card for Dataset Name
The Yelp reviews full star dataset is constructed by randomly taking 130,000 training samples and 10,000 testing samples for each review star from 1 to 5. In total there are 650,000 trainig samples and 50,000 testing samples.
Dataset Description
The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 2 columns in them, corresponding to class index (1 to 5) and review text. The review texts are… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Yelp_Reviews_for_Sentiment_Analysis_fine_grained_5_classes.student-depression-analysis
Assignment #1: EDA & Dataset
Predicting and Preventing Student Depression
Student: Amit GoodmanProgram: Economics & Entrepreneurship, Reichman University (RUNI)Date: March 2026
Project Overview
In this project, I explore the "Student Depression Dataset" to build a narrative around student well-being. By analyzing academic pressure, financial stress, and lifestyle habits, I aim to identify predictable risk factors and uncover actionable protective… See the full description on the dataset page: https://huggingface.co/datasets/ashlyannjophd2112/student-depression-analysis.Yelp_Reviews_for_Binary_Senti_Analysis
Dataset Card for Dataset Name
The Yelp reviews polarity dataset is constructed by considering stars 1 and 2 negative, and 3 and 4 positive. For each polarity 280,000 training samples and 19,000 testing samples are take randomly. In total there are 560,000 trainig samples and 38,000 testing samples. Negative polarity is class 1, and positive class 2.
Dataset Description
The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 2… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Yelp_Reviews_for_Binary_Senti_Analysis.
