datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-PolyBench
SWE-PolyBench
SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in each language is:
Javascript: 1017
Typescript: 729
Python: 199
Java: 165
Datasets
There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances and… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SWE-PolyBench.migration-bench-java-full
MigrationBench
1. 📖 Overview
🤗 MigrationBench
is a large-scale code migration benchmark dataset at the repository level,
across multiple programming languages.
Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.SWE-PolyBench_Verified
SWE-PolyBench
SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in the verified split is:
Javascript: 100
Typescript: 100
Python: 113
Java: 69
Datasets
There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SWE-PolyBench_Verified.SWE-PolyBench_500
SWE-PolyBench
SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in each language is:
Javascript: 1017
Typescript: 729
Python: 199
Java: 165
Datasets
There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances and… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SWE-PolyBench_500.sop-bench
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
📄 Paper: SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
🏭 Human Expert-Authored SOPs · 🤖 Human-AI Collaborative Framework · 📊 Executable Interfaces · 🔧 Two Agent Architectures · 📈 11 Frontier Models Evaluated
Dataset Summary
SOP-Bench is a comprehensive benchmark for evaluating LLM-based agents on complex, multi-step Standard Operating Procedures (SOPs) that are fundamental to industrial… See the full description on the dataset page: https://huggingface.co/datasets/amazon/sop-bench.amazon_review_fullmigration-bench-java-selected
MigrationBench
1. 📖 Overview
🤗 MigrationBench
is a large-scale code migration benchmark dataset at the repository level,
across multiple programming languages.
Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-selected.Quality-Control-App-Amazon-Big-Data-2023amazon-reviews-sentiment-analysis
Dataset Card for amazon reviews for sentiment analysis
Dataset Summary
One of the most important problems in e-commerce is the correct calculation of the points given to after-sales products. The solution to this problem is to provide greater customer satisfaction for the e-commerce site, product prominence for sellers, and a seamless shopping experience for buyers. Another problem is the correct ordering of the comments given to the products. The prominence of misleading… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/amazon-reviews-sentiment-analysis.amazon_zhthis is a datasets about amazon reviews
amazon_product_reviews_video_games#Title
amazon-fine-food-reviewsAmazon-C4
Amazon-C4
A complex product search dataset built based on Amazon Reviews 2023 dataset.
C4 is short for Complex Contexts Created by ChatGPT.
Quick Start
Loading Queries
from datasets import load_dataset
dataset = load_dataset('McAuley-Lab/Amazon-C4')['test']
>>> dataset
Dataset({
features: ['qid', 'query', 'item_id', 'user_id', 'ori_rating', 'ori_review'],
num_rows: 21223
})
>>> dataset[288]
{'qid': 288, 'query': 'I need something that can entertain my… See the full description on the dataset page: https://huggingface.co/datasets/McAuley-Lab/Amazon-C4.amazon-beauty-reviews-dataset
Dataset Card for "Amazon Beauty Reviews"
Dataset Summary
This dataset consists of reviews of "All Beauty" category from amazon. The data includes all ~700,000 reviews up to 2023. Reviews include product and user information, ratings, and a plain text review.
Supported Tasks and Leaderboards
This dataset can be used for numerous tasks like sentiment analysis, text classification, and user behavior analysis. It's particularly useful for training models to… See the full description on the dataset page: https://huggingface.co/datasets/jhan21/amazon-beauty-reviews-dataset.Amazon_Reviews_Binary_for_Sentiment_Analysis
Dataset Card for Dataset Name
The Amazon reviews polarity dataset is constructed by taking review score 1 and 2 as negative, and 4 and 5 as positive. Samples of score 3 is ignored. In the dataset, class 1 is the negative and class 2 is the positive. Each class has 1,800,000 training samples and 200,000 testing samples.
Dataset Details
Dataset Description
The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 3… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Amazon_Reviews_Binary_for_Sentiment_Analysis.migration-bench-java-utg
MigrationBench
1. 📖 Overview
🤗 MigrationBench
is a large-scale code migration benchmark dataset at the repository level,
across multiple programming languages.
Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-utg.amazon-stock-dataDescription of AMZN.csv
This CSV file contains historical stock price data for Amazon (AMZN) from May 15, 1997, to April 5, 2023.
The dataset includes the following columns:
Date: The date of the trading day.
Open: The opening price of the stock on that day.
High: The highest price the stock reached during the trading day.
Low: The lowest price the stock reached during the trading day.
Close: The closing price of the stock on that day.
Adj Close: The adjusted closing price, which accounts for… See the full description on the dataset page: https://huggingface.co/datasets/asyraffucl/amazon-stock-data.amazon_zh_simpleAmazon_Reviews_for_Sentiment_Analysis_fine_grained_5_classes
Dataset Card for Dataset Name
The Amazon reviews full score dataset is constructed by randomly taking 600,000 training samples and 130,000 testing samples for each review score from 1 to 5. In total there are 3,000,000 trainig samples and 650,000 testing samples.
Dataset Details
Dataset Description
The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 3 columns in them, corresponding to class index (1 to 5)… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Amazon_Reviews_for_Sentiment_Analysis_fine_grained_5_classes.zapolskii-amazondataset from kaggle https://www.kaggle.com/c/amazon-pet-product-reviews-classification
Multi-IaC-Eval
Multi-IaC-Eval
We present Multi-IaC-Eval is a novel benchmark dataset for evaluating LLM-based IaC generation and mutation across AWS CloudFormation, Terraform, and Cloud Development Kit (CDK) formats. The dataset consists of triplets containing initial IaC templates, natural language modification requests, and corresponding updated templates, created through a synthetic data generation pipeline with rigorous validation.
Cloudformation: 263
Terraform: 446
CDK (Python): 64
CDK… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/Multi-IaC-Eval.Amazon_2023_itemsxtr-wiki_qa
Xtr-WikiQA
Dataset Summary
Xtr-WikiQA is an Answer Sentence Selection (AS2) dataset in 9 non-English languages, proposed in our paper accepted at ACL 2023 (Findings): Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages.
This dataset is based on an English AS2 dataset, WikiQA (Original, Hugging Face).
For translations, we used Amazon Translate.
Languages
Arabic (ar)
Spanish (es)
French (fr)
German (de)
Hindi (hi)… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/xtr-wiki_qa.amazon_product_descriptionamazon-food-reviews-dataset
Dataset Card for "Amazon Food Reviews"
Dataset Summary
This dataset consists of reviews of fine foods from amazon. The data span a period of more than 10 years, including all ~500,000 reviews up to October 2012. Reviews include product and user information, ratings, and a plain text review. It also includes reviews from all other Amazon categories.
Supported Tasks and Leaderboards
This dataset can be used for numerous tasks like sentiment analysis, text… See the full description on the dataset page: https://huggingface.co/datasets/jhan21/amazon-food-reviews-dataset.amazon-us-fashion-deal-index
🛍️ Amazon US Fashion Deal Intelligence & Price Index (WFPI)
Official longitudinal time-series open dataset tracking verified Amazon US fashion pricing drops, category-level discount distributions, and brand clearance liquidations across 60,000+ fashion products. Published weekly by TheFashionDeals.com.
[!NOTE]
🌐 Dataset Scope & Hierarchy Breakdown
To maintain pristine data quality and prevent noisy scrapers, this repository is architected into two layers:… See the full description on the dataset page: https://huggingface.co/datasets/thefashiondeals/amazon-us-fashion-deal-index.amazon-product-data-2020
What is this?
This is a cleaned version of Amazon Product Dataset 2020 from Kaggle.
Why?
Using via Hugging Face API is easier; Kaggle API is annoying because their authentication is having credentials in a folder.
Cleaned because 13/28 columns are empty.
amazon_polarityamazon-reviews
Amazon Reviews Dataset
This dataset contains Amazon product reviews with binary sentiment labels (positive, negative) for text classification tasks.
Dataset Description
The dataset includes:
train.csv - Training set, 5000 samples
test.csv - Test set, 1000 samples
Usage
import pandas as pd
from huggingface_hub import hf_hub_download
# Download the training set
file_path = hf_hub_download(
repo_id="Cleanlab/amazon-reviews",
filename="train.csv"… See the full description on the dataset page: https://huggingface.co/datasets/Cleanlab/amazon-reviews.book-reviews-from-amazon-and-goodreads
Dataset Details
Dataset Description
Curated by: Lkkash
Language(s) (NLP): English
License: Apache license 2.0
Dataset Sources [optional]
Repository: https://huggingface.co/datasets/Lkkash/book-reviews-from-amazon-and-goodreads
Uses
Can be used to fine-tune text classification models directly from this dataset.
This dataset is used to train this model :- https://huggingface.co/Lkkash/distilbert-book-reviews
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Lkkash/book-reviews-from-amazon-and-goodreads.
