Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01evalitahf /sentiment_analysisSENTIPOLC 2016 dataset The SENTIPOLC 2016 dataset contains 9410 tweets annotated for subjectivity, overall and literal polarity, and irony. The dataset has been created and used in the context of the SENTIPOLC 2016 task (http://www.di.unito.it/~tutreeb/sentipolc-evalita16/index.html), organized as part of the EVALITA 2016 evaluation campaign. Original files available here: https://live.european-language-grid.eu/catalogue/corpus/7479/download/ If you find this dataset useful please cite:… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/sentiment_analysis.texttext-classification1K<n<10K0 likes513 downloads2y agoHugging Face02deep-analysis-research /simple-evalstext100K<n<1M0 likes443 downloads11mo agoHugging Face03GuarinoIndustriesLLC /guarino-dimensional-analysis-corpus Guarino Dimensional Analysis Corpus A machine-readable index of the published research programme of Brian Guarino (Guarino Industries LLC, ORCID 0009-0008-6705-8705). The programme applies one method, Buckingham π dimensional analysis, across domains that normally have nothing to do with each other: military effectiveness, defense sustainment, manufacturing, reactor thermal hydraulics, interdependent infrastructure, nurse staffing, consumer purchase decisions, the philosophy of… See the full description on the dataset page: https://huggingface.co/datasets/GuarinoIndustriesLLC/guarino-dimensional-analysis-corpus.textn<1K0 likes397 downloads7h agoHugging Face04daaain /swebench-verified-deepseek-v4-flash-failure-analysis SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model driven by mini-swe-agent, graded with the official SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative root-cause diagnosis. Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.tabulartext-generationn<1K0 likes306 downloads4mo agoHugging Face05RobinChen2001 /A-Survey-for-LLM-Agent-Trajectory-Analysis A Survey for LLM Agent Trajectory Analysis This dataset repository hosts the survey paper A Survey for LLM Agent Trajectory Analysis: From Failure Attribution to Enhancement and a structured metadata snapshot of the companion paper collection from Awesome-LLM-Agent-Trajectory-Analysis. The repository is intended for discovery, citation, and lightweight analysis of the literature around LLM agent trajectory analysis, including failure attribution, trajectory-based debugging… See the full description on the dataset page: https://huggingface.co/datasets/RobinChen2001/A-Survey-for-LLM-Agent-Trajectory-Analysis.documentn<1K2 likes222 downloads3mo agoHugging Face06zjunlp /DataMind-Analysis-SFT-DataThis repository contains the data presented in Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study Code: https://github.com/zjunlp/DataMind text1K<n<10K1 likes172 downloads1y agoHugging Face07khaihernlow /massive-stock-news-analysis-db-for-nlpbackteststext1M<n<10M0 likes149 downloads2y agoHugging Face08openerotica /erotica-analysisThis dataset is roughly 27k examples of erotica stories which I've fed through GPT-3.5-turbo-16k to obtain a summary, writing prompt, and tags as a response. I've filtered out all the refusals, and deleted a fair ammount of "GPT-isms". I'd still like to go through this again to prune any remaining low quality responses I've missed, but I think this is a good start. Most of the context size comes from the stories themselves, not the responses. Please consider supporting my Patreon… See the full description on the dataset page: https://huggingface.co/datasets/openerotica/erotica-analysis.text10K<n<100K36 likes148 downloads2y agoHugging Face09mehti /LMOD-Cataract-1K-surgical-analysis-cot Cataract-1K LLM-Generated Surgical Instructions Dataset Overview This dataset is derived from the Cataract-1K dataset (part of the LMOD benchmark) and enhanced using Qwen3-VL-30B-A3B-Thinking, a large vision-language model with reasoning capabilities. It is designed for training medical AI systems to provide actionable surgical guidance with transparent reasoning. Generation Process Source Data: Cataract-1K processed frames with segmentation annotations… See the full description on the dataset page: https://huggingface.co/datasets/mehti/LMOD-Cataract-1K-surgical-analysis-cot.imagevisual-question-answering10K<n<100K0 likes148 downloads8mo agoHugging Face10RegalFire /Scientific-Code-and-Analysis-QA RegalFire Scientific code and analysis QA RegalFire — AI Data Foundry Scientific forum QA containing mechanically extracted preformatted code. Code is retained exactly after HTML entity decoding; no execution is claimed. Verified scope Records: 72; distinct source threads: 72; unique answers represented: 99. Domain thread counts: {"statistics": 34, "computational_science": 37, "biology": 1}. Actual record splits: {"train": 48, "holdout": 12, "test": 6… See the full description on the dataset page: https://huggingface.co/datasets/RegalFire/Scientific-Code-and-Analysis-QA.textn<1K0 likes138 downloads10d agoHugging Face11max-rl /perplexity_analysis Perplexity Analysis This repository contains the data, scripts, and generated figures used for perplexity analysis experiments. Contents data/Qwen3: rollout data for Qwen3 1.7B and 4B Base, GRPO, and MaxRL models on AIME25 and BeyondAIME. data/Maze/perplexity: maze rollout data and derived perplexity analysis artifacts. outputs: generated JSON summaries and figures for Qwen3 analyses. *.py: analysis and plotting scripts. See data/README.md for additional data details… See the full description on the dataset page: https://huggingface.co/datasets/max-rl/perplexity_analysis.image10M<n<100M0 likes109 downloads6mo agoHugging Face12yuncongli /chat-sentiment-analysis A Sentiment Analsysis Dataset for Finetuning Large Models in Chat-style More details can be found at https://github.com/l294265421/chat-sentiment-analysis Supported Tasks Aspect Term Extraction (ATE) Opinion Term Extraction (OTE) Aspect Term-Opinion Term Pair Extraction (AOPE) Aspect term, Sentiment, Opinion term Triplet Extraction (ASOTE) Aspect Category Detection (ACD) Aspect Category-Sentiment Pair Extraction (ACSA) Aspect-Category-Opinion-Sentiment (ACOS) Quadruple… See the full description on the dataset page: https://huggingface.co/datasets/yuncongli/chat-sentiment-analysis.text10K<n<100K9 likes88 downloads4y agoHugging Face13MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-anger Dataset Summary Synthetic Persian Chatbot Conversational SA – Anger is a Persian (Farsi) dataset created for the Classification task, with a focus on detecting the emotion "anger" in chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is derived from the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset. Language(s): Persian (Farsi) Task(s): Classification… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-anger.text1K<n<10K0 likes88 downloads1y agoHugging Face14MCINext /chatbot-conversational-sentiment-analysis-tone-user-classification Dataset Summary Synthetic Persian Chatbot Conversational SA – User Tone Classification(SynPerChatbotConvSAToneUserClassification) is a Persian (Farsi) dataset created for the Classification task. It focuses on identifying the user’s conversational tone—formal, casual, or childish—in emotionally rich chatbot interactions. This dataset is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using GPT-4o-mini. Language(s): Persian (Farsi)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/chatbot-conversational-sentiment-analysis-tone-user-classification.text1K<n<10K0 likes88 downloads1y agoHugging Face15MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-friendship Dataset Summary Synthetic Persian Chatbot Conversational SA – Friendship is a Persian (Farsi) dataset created for the Classification task, with a focus on detecting the emotion "friendship" in chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is derived from the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset. Language(s): Persian (Farsi) Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-friendship.textn<1K0 likes83 downloads1y agoHugging Face16MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-fear Dataset Summary Synthetic Persian Chatbot Conversational SA – Fear is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "fear" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is a subset of the Synthetic Persian Chatbot Conversational Sentiment Analysis dataset. Language(s): Persian (Farsi) Task(s): Classification (Emotion… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-fear.textn<1K0 likes83 downloads1y agoHugging Face17MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-sadness Dataset Summary Synthetic Persian Chatbot Conversational SA – Sadness is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "sadness" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset. Language(s): Persian (Farsi) Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-sadness.textn<1K0 likes80 downloads1y agoHugging Face18MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-satisfaction Dataset Summary Synthetic Persian Chatbot Conversational Sentiment Analysis – Satisfaction is a Persian (Farsi) dataset developed for the Classification task, specifically focused on detecting the emotion of satisfaction in chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using the GPT-4o-mini language model. Language(s): Persian (Farsi) Task(s): Classification (Emotion Detection – Satisfaction) Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-satisfaction.text1K<n<10K0 likes78 downloads1y agoHugging Face19MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-jealousy Dataset Summary Synthetic Persian Chatbot Conversational SA – Jealousy is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "jealousy" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis dataset. Language(s): Persian (Farsi) Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-jealousy.textn<1K0 likes78 downloads1y agoHugging Face20MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-surprise Dataset Summary Synthetic Persian Chatbot Conversational SA – Surprise is a Persian (Farsi) dataset for the Classification task, focused on detecting the expression of "surprise" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is a subset of the Synthetic Persian Chatbot Conversational Sentiment Analysis dataset. Language(s): Persian (Farsi) Task(s): Classification… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-surprise.textn<1K0 likes75 downloads1y agoHugging Face21MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-love Dataset Summary Synthetic Persian Chatbot Conversational SA – Love is a Persian (Farsi) dataset for the Classification task, specifically focused on detecting the emotion "love" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis collection. Language(s): Persian (Farsi) Task(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-love.textn<1K0 likes74 downloads1y agoHugging Face22MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-happiness Dataset Summary Synthetic Persian Chatbot Conversational SA – Happiness is a Persian (Farsi) dataset for the Classification task, specifically focused on detecting the emotion "happiness" in user-chatbot conversations. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using GPT-4o-mini, and is a subset of the broader Synthetic Persian Chatbot Conversational Sentiment Analysis collection. Language(s): Persian (Farsi)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-happiness.texttext-classificationn<1K0 likes73 downloads1y agoHugging Face23max-rl /variance_analysis Variance Analysis This repository contains the data, scripts, and generated figures used for variance analysis experiments. Contents data/SmolLM: SmolLM GSM8K rollout data. data/Qwen3: Qwen3 math rollout data, including 1.7B 512 x 512 and 4B 1024 x 128 samples. data/Maze/variance: Maze rollout data for variance analysis. outputs: generated JSON summaries and figures. *.py and run_*.sh: analysis, plotting, and Slurm launch scripts. See data/README.md for additional data… See the full description on the dataset page: https://huggingface.co/datasets/max-rl/variance_analysis.document1M<n<10M0 likes73 downloads6mo agoHugging Face24MCINext /synthetic-persian-chatbot-conversational-sentiment-analysis-tone-chatbot-classification Dataset Summary Synthetic Persian Chatbot Conversational SA – Chatbot Tone Classification(SynPerChatbotConvSAToneChatbotClassification) is a Persian (Farsi) dataset created for the Classification task. It focuses on identifying the chatbot’s conversational tone—formal, casual, or childish—in dialogue exchanges that include emotional content. This dataset is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using GPT-4o-mini. Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-conversational-sentiment-analysis-tone-chatbot-classification.text1K<n<10K0 likes72 downloads1y agoHugging Face25Daniel-ML /sentiment-analysis-for-financial-news-v2text1K<n<10K1 likes70 downloads2y agoHugging Face26OdiaGenAI /sentiment_analysis_hindiConventions followed to decide the polarity: - labels consisting of a single value are left undisturbed, i.e. if label = 'pos', then it'll be pos labels consisting of multiple values separated by '&' are processed. If all the labels are the same ('pos&pos&pos' or 'neg&neg'), then the shortened form of the multiple label is assigned as the final label. For example, if label = 'pos&pos&pos', then final label will be 'pos'. labels consisting of mixed values ('pos&neg&pos' or 'neg&neu&pos') are… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/sentiment_analysis_hindi.texttext-classification1K<n<10K2 likes68 downloads3y agoHugging Face27NNEngine /Sentiment-Analysis-ComplexExcellent — congrats on getting the repo ready 🚀 Here’s a professional Hugging Face Dataset Card (README.md) you can paste directly into your repository. This is written to match HF best practices and serious research usage. 📘 README.md 👉 Copy everything below into your README.md Sentiment-Analysis-Complex 🧠 Overview Sentiment-Analysis-Complex is a large-scale synthetic sentiment analysis dataset designed for benchmarking modern NLP models under… See the full description on the dataset page: https://huggingface.co/datasets/NNEngine/Sentiment-Analysis-Complex.texttext-classification10M<n<100M0 likes68 downloads9mo agoHugging Face28bdstar /twitter-sentiment-analysis 🐦 Twitter Sentiment Analysis (bdstar/twitter-sentiment-analysis) 🧠 Overview A refined and merged version of Twitter text sentiment datasets, providing a clean and well-balanced dataset for sentiment classification across three sentiment categories:positive, negative, and neutral. This dataset is split into three parts — train, test, and validation — each sourced from highly reputable open datasets.It is designed for training, evaluating, and benchmarking NLP models for… See the full description on the dataset page: https://huggingface.co/datasets/bdstar/twitter-sentiment-analysis.texttext-classification1M<n<10M0 likes67 downloads1y agoHugging Face29saillab /medical-vqa-robustness-analysis Medical VQA Robustness Analysis This dataset contains robustness analysis results for medical vision-language models on chest X-ray visual question answering tasks. The analysis evaluates model performance under various question perturbations to assess clinical safety and response stability. Source Data This analysis is based on the MIMIC-CXR-VQA dataset from PhysioNet, which provides chest X-ray images paired with clinically relevant questions and answers. Models… See the full description on the dataset page: https://huggingface.co/datasets/saillab/medical-vqa-robustness-analysis.textvisual-question-answeringn<1K0 likes62 downloads1y agoHugging Face30newsbang /math_benbench_data_leak_analysis Dataset description This is a math dataset mixed from four open-source data. It was used to analyze the contamination test on the MATH and contains 1M samples. Dataset fields question    question from open-source data solution    the answer corresponding to question 5grams    5-gram list of f"{question} {answer}" test_question    the most relevant question from MATH test_solution    the answer corresponding to test_question test_5grams    5-gram list of… See the full description on the dataset page: https://huggingface.co/datasets/newsbang/math_benbench_data_leak_analysis.text100K<n<1M3 likes61 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.