datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
data_analysis
Dataset Card for "livebench/data_analysis"
LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties:
LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses.
Each question has verifiable, objective ground-truth answers, allowing hard questions to be… See the full description on the dataset page: https://huggingface.co/datasets/livebench/data_analysis.DataMind-Analysis-SFT-DataThis repository contains the data presented in Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study
Code: https://github.com/zjunlp/DataMind
sentiment_analysis_data
Dataset Card for "sentiment_analysis_data"
More Information needed
open-pulse-hackathon-data-analysis
LauzHack Projects Dataset
Dataset Summary
This dataset contains comprehensive information about projects submitted to
LauzHack (EPFL's student-run hackathon) from 2023 to 2025. Each project
includes details about the project title, description, team members, awards, and
categories.
LauzHack is an annual 24-hour hackathon hosted at EPFL (École Polytechnique
Fédérale de Lausanne) in Lausanne, Switzerland, bringing together students and
hackers to create innovative solutions… See the full description on the dataset page: https://huggingface.co/datasets/SDSC/open-pulse-hackathon-data-analysis.code_completion_for_data_analysisdata-table-analysisThis dataset can be used for benchmarking LLM Structured Outputs via the code here:
https://github.com/cleanlab/structured-output-benchmark/
math_benbench_data_leak_analysis
Dataset description
This is a math dataset mixed from four open-source data. It was used to analyze the contamination test on the MATH and contains 1M samples.
Dataset fields
question
question from open-source data
solution
the answer corresponding to question
5grams
5-gram list of f"{question} {answer}"
test_question
the most relevant question from MATH
test_solution
the answer corresponding to test_question
test_5grams
5-gram list of… See the full description on the dataset page: https://huggingface.co/datasets/newsbang/math_benbench_data_leak_analysis.text-analysis-context-cased-case-dataNostalgic_Sentiment_Analysis_of_YouTube_Comments_Data
Dataset Summary
The dataset is a collection of Youtube Comments and it was captured using the YouTube Data API.
The data set consists of 1500 nostalgic and non-nostalgic comments in English.
Languages
The language of the data is English.
Citation
If you find this dataset usefull for your study, please cite the paper as followed:
@article{postalcioglu2020comparison,
title={Comparison of Neural Network Models for Nostalgic Sentiment Analysis of YouTube… See the full description on the dataset page: https://huggingface.co/datasets/Senem/Nostalgic_Sentiment_Analysis_of_YouTube_Comments_Data.data-analysis-datasets
CANNS Analysis Datasets
This repository contains example datasets for the CANNS (Continuous Attractor Neural Networks) data analysis package.
Datasets
ROI_data.txt (703 KB)
Description: 1D CANN ROI data for bump analysis
Format: Text file with neural activity measurements
Usage: 1D CANN analysis, MCMC bump fitting
Example: Used in 1D CANN analysis tutorials
grid_1.npz (8.7 MB)
Description: Grid cell spike data with position information
Format:… See the full description on the dataset page: https://huggingface.co/datasets/canns-team/data-analysis-datasets.sentiment_analysis_financial_news_datalivebench-data_analysis
Dataset Card for "livebench/data_analysis"
LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties:
LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses.
Each question has verifiable, objective ground-truth answers, allowing hard questions to be… See the full description on the dataset page: https://huggingface.co/datasets/lthn/livebench-data_analysis.sentiment-analysis-for-mental-health-Combined-DataData_Analysis_Workflow_20240610_100454Data_Analysis_Workflow_20240509_134737matplotlib_data_analysis_questionsData_Analysis_Workflow_20240425_111154Data_Analysis_Workflow_20240509_121811Twitter-mulitlingual-synthetic-data-sentiment-analysis
Dataset Card for Twitter-mulitlingual-synthetic-data-sentiment-analysis
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/Paul-HF/Twitter-mulitlingual-synthetic-data-sentiment-analysis/raw/main/pipeline.yaml"
or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/Paul-HF/Twitter-mulitlingual-synthetic-data-sentiment-analysis.data-analysis-ai-agent
Data Analysis Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP
This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc.
The dataset is helpful for AI… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/data-analysis-ai-agent.Data_Analysis_Workflow_20240504_054612drone_turn_analysis_data
Dataset Summary
Drone Turn Analysis Data dataset provides information on drone turning behavior, this dataset collected using the Mission Planner simulator for educational purpose.
Features:
Timestamp: Time of the data capture.
Velocity: Drone's speed (m/s) during the turn.
Heading Change: Directional change in degrees.
Target Distance: Planned distance for the maneuver.
Actual Distance: Actual distance covered.
Overshoot: Difference between target and actual distances.… See the full description on the dataset page: https://huggingface.co/datasets/SGF-14/drone_turn_analysis_data.NCSS_2023_Data_Analysisinstruct_code_for_data_analysisData_Analysis_Workflow_20240503_224446nile-gcs-processed-qa-yolo-scaled-up-verifiable-data-analysisEncoding_Mismatch_Analysis_Data
Encoding Mismatch Analysis Data
This repository publishes the prepared numerical analysis artifacts associated
with From Per-Image Low-Rank to Encoding Mismatch: Rethinking Feature
Distillation in Vision Transformers. It is analysis data, not an image or
model-training dataset, and it does not redistribute ImageNet.
Links
Paper: https://arxiv.org/abs/2511.15572
Hugging Face paper page: https://huggingface.co/papers/2511.15572
Code and analysis scripts:… See the full description on the dataset page: https://huggingface.co/datasets/Huiyuancs/Encoding_Mismatch_Analysis_Data.1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample
Description
Korean Test Questions Structured Analysis Processing Data, around 1.5 million questions, contains question types, questions, answers, explanations, etc..For subjects, include [Primary School] Korean, Mathematics, English, Social Studies, Science; [Middle School] Korean, English, Mathematics, Science, Social Studies; [High School] Korean, English, Mathematics, Physics, Chemistry, Biology, History, Geography; question Types indlude single-choice question, fill-in… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.Data_Analysis_Workflow_20240503_195731Data_Analysis_Workflow_20240504_054711
