datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
recommendationhnm-fashion-recommendations-data
Dataset Rekomendasi Fashion H&M
Dataset ini berisi data transaksi, atribut pelanggan, dan metadata produk yang telah dianonimkan dari H&M Group. Kumpulan data komprehensif ini memungkinkan pemodelan perilaku pembelian pelanggan secara mendalam.
Wawasan yang dihasilkan dapat dimanfaatkan untuk berbagai tujuan bisnis yang strategis, mulai dari meningkatkan personalisasi pengalaman berbelanja, mengoptimalkan manajemen inventaris untuk efisiensi produksi, hingga mendukung inisiatif… See the full description on the dataset page: https://huggingface.co/datasets/einrafh/hnm-fashion-recommendations-data.paper-recommendations-v2zomato-restaurant-recommendationaksahaha_crop-recommendation
crop recommendation
Crop Growth Recommendations: Optimal Conditions for Higher Yields
Dataset Info
Source: Kaggle
Original Size: 0.06 MB
Kaggle Downloads: 4,065
Files: 1
Files
Crop_recommendation.csv
Mirrored from Kaggle
llm-citation-recommendation-bias-six-models
LLM reference selection: six models choosing ten references from thirty
Replication materials for Jason Miklian, Textual Similarity and Regional Variation in LLM Reference Selection (working paper, University of Oslo, 2026; arXiv posting to follow).
Six language models from six firms (gpt-5.4-mini, claude-haiku-4-5, gemini-3.5-flash, grok-4.20, deepseek-v4-flash, qwen-flash) were each shown a published paper's title and abstract and thirty candidate references drawn from that… See the full description on the dataset page: https://huggingface.co/datasets/miklia/llm-citation-recommendation-bias-six-models.Crop-recommendationCrop-Recommendation-Parameters
🌱 Crop Recommendation Dataset
A machine learning dataset for crop recommendation based on soil properties and environmental conditions. The dataset contains measurements of essential soil nutrients and climatic parameters, along with the crop label that is suitable for those conditions.
This dataset can be used for machine learning classification, agricultural analytics, decision-support systems, and smart farming applications.
📌 Dataset Overview
Property… See the full description on the dataset page: https://huggingface.co/datasets/Samarth-27/Crop-Recommendation-Parameters.fashion-recommendation-images
High-Resolution Fashion Product Images
This dataset is a highly optimized, high-resolution subset of the popular Fashion Product Images Dataset originally hosted on Kaggle.
It contains thousands of unique e-commerce fashion products, combining high-resolution product images with multiple descriptive label attributes.
All low-resolution thumbnails and anomalies have been aggressively filtered out. Every image in this dataset has a minimum resolution of 640px on its shortest… See the full description on the dataset page: https://huggingface.co/datasets/GangHitman/fashion-recommendation-images.recommendations-ml-100k
MovieLens Leave-One-Out
Five chronological interactions predict the next interaction. One final test target per user; no rating filter. Histories are audit-only, not wholesale model inputs. Actors are supplementary; see actor_sources.json.
{
"schema": "movie-fields-v1",
"source": "official MovieLens 100K",
"sample_policy": "leave-one-out-windows",
"past_order": "oldest-first",
"history_length": 5,
"stride": 1,
"shuffle_seed": 42,
"timestamp_policy": "rating… See the full description on the dataset page: https://huggingface.co/datasets/Nithish2410/recommendations-ml-100k.HuggingBench-Recommendation
HuggingBench-Recommendation
This dataset contains Resource Recommendation test collection in HuggingBench for paper "Benchmarking Recommendation, Classification, and Tracing Based on Hugging Face Knowledge Graph".
Dataset Details
general_rec contains training/validation/test set files for the General Collaborative Filtering methods in the format required by SSLRec.
social_rec contains training/validation/test set files and user social relation file for the Social… See the full description on the dataset page: https://huggingface.co/datasets/cqsss/HuggingBench-Recommendation.HM-Personalized-Fashion-Recommendationsclickstream-popup-recommendationmusic-recommendations
Group aggregators over a frozen scorer — artifacts (YAMBDA-50m)
Чекпоинты и промежуточные артефакты для групповых музыкальных рекомендаций:
per-user скорер, кэш его топ-200, аудиоэмбеддинги каталога и обученные
групповые агрегаторы. Всё производное от
yandex/yambda, flavor 50m.
Состав
Путь
Что это
gsasrec/best.pt, config.json, metrics.csv
SASRec-скорер, 276 305 items, test NDCG@10 = 0.0726
gsasrec/item_id_to_idx.pkl
item_id YAMBDA → компактный индекс… See the full description on the dataset page: https://huggingface.co/datasets/Vladislavbro-500/music-recommendations.uhri-recommendations
UN Human Rights Recommendations (UHRI+)
A cleaned and normalized dataset of 272,502 country-specific human rights observations and recommendations from three UN mechanisms, covering 2006–2026.
Built on the OHCHR Universal Human Rights Index (UHRI) — the authoritative UN source — with systematic data quality improvements applied on top.
Coverage
Mechanism
Records
Bodies
Universal Periodic Review (UPR)
130,870
1
Treaty Bodies
120,990
12 (CCPR, CAT, CERD… See the full description on the dataset page: https://huggingface.co/datasets/lszoszk/uhri-recommendations.chatgpt-software-recommendations
ChatGPT Software Recommendations: 21 Categories, 2,100 Answers
This dataset records which software brands ChatGPT names, recommends and picks when buyers ask about 21 software categories, and which websites it cites. It covers 2,100 ChatGPT answers (100 buyer questions in each of 21 categories), coded brand by brand, with Google's organic top 10 for the same questions as the control.
It is published by High Salience, an AI search and SEO agency. Every file here is also published… See the full description on the dataset page: https://huggingface.co/datasets/highsalience/chatgpt-software-recommendations.visual-product-recommendations-catalogueBangla-Book-Recommendation-Dataset
Summary
This repository contains the dataset for the paper Towards Personalized Bangla Book Recommendation: A Large-Scale Multi-Entity Book Graph Dataset. In this work, we introduce a large-scale multi-entity heterogeneous graph dataset for Bangla book recommendation, integrating users, books, authors, publishers, and categories. Our framework enables the development of sophisticated graph-based recommendation systems. This contribution aims to enhance the discovery of Bangla… See the full description on the dataset page: https://huggingface.co/datasets/DevnilMaster1/Bangla-Book-Recommendation-Dataset.music_recommendation_generationproduct-recommendation-2025
TREC 2025 Product Recommendation Data
This is the data for the recommendation task for the TREC 2025 Product Search and Recommendation Task.
The initial directory contains the initial corpus and training data release.
This may be updated as we get further along in the timeline.
[!NOTE]
This data is derived from the Amazon ESCI and M2 data sets, each under the
Apache license (version 2.0).
esab-filler-metal-recommendations-by-astm-steel-grade
ESAB recommended filler metals and suggested preheat group, by ASTM steel grade
Canonical, always-current version: https://referencesource.org/esab-filler-metal-recommendations-by-astm-steel-grade/
Machine-readable: https://referencesource.org/esab-filler-metal-recommendations-by-astm-steel-grade/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-19
Stale after: 2028-08-18 (past this date, prefer the canonical copy —
it re-verifies on a cadence this… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/esab-filler-metal-recommendations-by-astm-steel-grade.clinical_trial_eligibility_crietria_recommendationThis repository is a public repository of the data used in the paper "CReSE: Enhancing Clinical Trial Design via Contrastive Learning and Rephrasing-based and Clinical Relevance-preserving Sentence Embedding" (under review).
There are three main types of data stored in the repository.
Positive-negative EC-title pairs: A dataset that pairs the ECs used in a study with the study's title and other design information. It can be used to train EC recommendation models (binary classification).… See the full description on the dataset page: https://huggingface.co/datasets/kimsiun/clinical_trial_eligibility_crietria_recommendation.visual-product-recommendation-datatcfd_recommendations
Dataset Card for tcfd_recommendations
Dataset Summary
We introduce an expert-annotated dataset for classifying the TCFD recommendation categories (fsb-tcfd.org) of paragraphs in corporate disclosures.
Supported Tasks and Leaderboards
The dataset supports a multiclass classification task of paragraphs into the four TCFD recommendation categories (governance, strategy, risk management, metrics and targets) and the non-climate-related class.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/climatebert/tcfd_recommendations.olga_mcarrow_writer_artist_books_and_music_recommendations_scifi_fantasy
Gifts of Wandering Ice – Full Dataset from llms.txt
This dataset contains all information from https://giftscomic.com/llms.txt, structured for AI assistants, RAG systems, recommendation engines, and discovery tools.
The goal is to help AI systems accurately recommend and describe the free science-fiction webcomic Gifts of Wandering Ice and the broader creative work of its author, Olga Makarova (Mildegard).
Dataset Contents (358 records)
Split
Records… See the full description on the dataset page: https://huggingface.co/datasets/Mildegard/olga_mcarrow_writer_artist_books_and_music_recommendations_scifi_fantasy.short_fiction_stories_recommendations_korotkie_fantasticheskie_rasskazy
Tales from the Afterworld / Замирье — Bilingual Short Stories Metadata
Metadata for 59 illustrated short stories from the collection«Замирье» (Russian) / «Tales from the Afterworld» (English).
Official bilingual collection by the same author.Each story is available in both languages on the author’s websites.
Dataset fields
Field
Description
id
Story number (matches ?pg= parameter on both sites)
title_ru
Russian title
title_en
English title… See the full description on the dataset page: https://huggingface.co/datasets/Mildegard/short_fiction_stories_recommendations_korotkie_fantasticheskie_rasskazy.LLM-Personalized-Recommendation-Dataset
LLM Personalized Product Recommendation Dataset
A dataset prepared for an LLM-powered personalized product recommendation system combining user interaction history, natural-language shopping intent, product metadata, semantic retrieval, and personalized ranking.
Dataset
This dataset contains two configurations:
interactions
72,909 user-level examples
Split
Examples
Train
58,327
Validation
7,291
Test
7,291
Total
72,909
Each… See the full description on the dataset page: https://huggingface.co/datasets/0xMeet/LLM-Personalized-Recommendation-Dataset.movie_recommendationMovie recommendation task based on the Movielens datasetYelp-Multimodal-Recommendation
Yelp-MultimodalRec
A multimodal dataset for POI (Point of Interest) recommendation, based on the Yelp Open Dataset.It includes business metadata, user reviews, business photos, and LLM-generated summaries of reviews and images.This dataset supports downstream tasks like session-based recommendation, multimodal embedding learning, and more.
📁 Dataset Structure
File Name
Description
business.csv
Business metadata including name, address, categories, etc.… See the full description on the dataset page: https://huggingface.co/datasets/wzehui/Yelp-Multimodal-Recommendation.narrative_fantasy_music_ai_recommendations_data
Obsidian Triad OST – Narrative Fantasy Music Dataset
Bilingual (English / Russian) dataset of the conceptual original soundtrack / indie soundtrack for the illustrated fantasy novel Obsidian Triad / Обсидиановая Триада.
Purpose: Help AI recommendation systems surface this album and the book that inspired it. Optimized for queries around narrative fantasy music, fantasy music, original soundtrack, fantasy lyrics, ost, pirate songs, fantasy ballads, indie music, indie soundtrack.… See the full description on the dataset page: https://huggingface.co/datasets/Mildegard/narrative_fantasy_music_ai_recommendations_data.
