datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
finepdfs_lang_classificationCobot_Magic_classification_of_tableware
Cobot_Magic_classification_of_tableware
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_tableware.Cobot_Magic_classification_of_fruits_and_vegetables
Cobot_Magic_classification_of_fruits_and_vegetables
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_fruits_and_vegetables.Cobot_Magic_classification_of_fruits_and_vegetables_a
Cobot_Magic_classification_of_fruits_and_vegetables_a
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_fruits_and_vegetables_a.mbti_classification_dataset_fullPosts
MBTI Classification Dataset (Full Posts)
A dataset of 8,675 personality-forum posts labeled with Myers-Briggs Type Indicator (MBTI) dichotomies, built for training and evaluating text-based personality classification models. Each row contains one user's concatenated forum posts plus binary labels for all four MBTI dimensions.
Dataset Structure
Splits: train (5,205 rows), test (2,082 rows), validation (1,388 rows)
Fields:
Field
Type
Description
I/E
int64… See the full description on the dataset page: https://huggingface.co/datasets/ClaudiaRichard/mbti_classification_dataset_fullPosts.source-classifications
NuBerea Source Gold Set
Curated source-critical classifications for the Hebrew Bible, New Testament, and Septuagint — the classical concerns of source criticism (documentary strata in the Old Testament, corpus structure in the New Testament, translation traditions in the Septuagint) expressed as structured, verse-level data, together with statistical validation summaries and characteristic-vocabulary ("hallmark") term lists.
This dataset is part of the NuBerea curated corpus… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/source-classifications.composition-classifications
NuBerea Composition Classifications
A curated reference set of scholarly-consensus composition history for the biblical corpus: the traditions behind the Old Testament, Deuterocanon, New Testament, and Old Testament Pseudepigrapha, and the source-critical relationships among them (e.g. Documentary Hypothesis strands, Markan priority, canonical collection, translation into the Septuagint). The dataset is a direct transcription of established scholarship — no machine learning or… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/composition-classifications.bi-so101-fruits-classificationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_so101_follower",
"total_episodes": 2,
"total_frames": 2910,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jy13/bi-so101-fruits-classification.CUAD_v1_Contract_Understanding_clause_classification
Dataset Card for Contract Understanding Atticus Dataset (CUAD) Clause Classification
This dataset contains 13,155 labeled clauses extracted from 509 commercial legal contracts from the original CUAD dataset. One of the original 510 contracts was removed due to being a scanned copy.
The text was cleaned using clean-text.
You can easily and quickly load it:
dataset = load_dataset("dvgodoy/CUAD_v1_Contract_Understanding_clause_classification")
Dataset({
features: ['file_name'… See the full description on the dataset page: https://huggingface.co/datasets/dvgodoy/CUAD_v1_Contract_Understanding_clause_classification.JDReview-classification
Dataset Card for "JDReview-classification"
More Information needed
methyl-classification
DNA Methylation Tissue Classification Dataset
Dataset Summary
Homepage: https://github.com/ylaboratory/methylation-classification
Pubmed: False
Public: True
This data resource is vast, curated reference atlas of DNA methylation (DNAm) profiles
spanning 16,959 healthy primary human tissue and cell samples profiled on Illumina 450K arrays.
Samples cover 86 unique tissues and cell types and are manually mapped to a common set of terms in the UBERON anatomical… See the full description on the dataset page: https://huggingface.co/datasets/ylab/methyl-classification.reddit-story-niche-classification-dataset
🧠 Reddit Niche Classification Dataset
This dataset contains 13,061 Reddit posts annotated with a custom niche label (e.g. advice, drama, humor, unknown, etc). It includes structured features engineered from post metadata, not raw text — making it ideal for lightweight classification models.
🧾 Schema
Column
Type
Description
title
string
Post title
selftext
string
Post body text
subreddit
string
Subreddit the post belongs to
flair
string
Flair… See the full description on the dataset page: https://huggingface.co/datasets/atin5551/reddit-story-niche-classification-dataset.ko-voicephishing-binary-classificationchatlabel-cn-id-shipping-classification
当前发布版本 v2026.09.30.2
本版对既有标注做跨文件来源对账、去重关系汇总和置信依据展示。当前商品保持 6921 条,历史 23414 条,图片 6179 张。没有重新标注、修改标签/备注或提高训练资格。旧 annotations 等5个配置的字节内容保留,因此其行内 dataset_version 仍表示原标注版本 2026.09.30.1;本次新增汇总配置版本为 2026.09.30.2。
新增 consolidation/current:每个当前 record_id 一条,包含全部来源、历史数量、关联商品、缺失字段、原始自报置信信息及来源。新增 duplicate_groups/groups:保守的同图、复用SKU、全证据相同但不同SKU关联组;不是自动确认的相同商品。4组/8条全证据相同但SKU不同的候选仍保留独立记录。
confidence_status 是审核状态,不是概率或准确率。模型自报值与旧自动化分数分开保留,全部未校准;不因重复文件增票,不把用户修订来源声明等同于逐条实名人审。详细方法见… See the full description on the dataset page: https://huggingface.co/datasets/Rex2wzh/chatlabel-cn-id-shipping-classification.Color_Obj_ClassificationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_shoulder_pan.pos",
"left_shoulder_lift.pos",
"left_elbow_flex.pos",
"left_wrist_flex.pos",
"left_wrist_yaw.pos",
"left_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/joonyoung-seong/Color_Obj_Classification.cybersec-topic-classification-dataset
Cybersecurity Topic Classification (CTC) Dataset
Note: This is an unofficial upload of the Cybersecurity Topic Classification (CTC) dataset. The original dataset and accompanying paper were developed by Elijah Pelofske, Lorie M. Liebrock, and Vincent Urias.
This dataset comprises training and validation data for the Cybersecurity Topic Classification (CTC) tool, as introduced in the paper "A Robust Cybersecurity Topic Classification Tool" by Elijah Pelofske, Lorie M. Liebrock, and… See the full description on the dataset page: https://huggingface.co/datasets/naufalso/cybersec-topic-classification-dataset.pyroptosis_classificationdata_ClassificationModel
Dataset Card for "data_ClassificationModel"---
dataset_info:
features:
- name: brands
dtype: string
- name: categories
dtype: string
- name: code
dtype: string
- name: languages_tags
dtype: string
- name: last_modified_t
dtype: int64
- name: product_name_de
dtype: string
- name: quantity
dtype: string
- name: index_level_0
dtype: int64
splits:
- name: train
num_bytes: 231023
num_examples: 673
download_size:… See the full description on the dataset page: https://huggingface.co/datasets/CedRuiz/data_ClassificationModel.stackoverflow-unified-text-open-status-classification
Dataset Card for "stackoverflow-unified-text-open-status-classification"
More Information needed
GTZAN-Dataset-Music-Genre-Classification
Hướng dẫn sử dụng Dataset GTZAN cho Model Team
1. Thành phần bàn giao
Dataset trên Hugging Face:
Tài nguyên đi kèm: [stats.json] , [label_map.json]
2. Cách Load Dataset từ Hugging Face
Dữ liệu đã chia sẵn thành 3 tập: train, validation và test theo tỷ lệ chuẩn, đảm bảo Zero-Leakage (các đoạn cắt từ cùng một bài hát gốc sẽ nằm chung trong một tập).
from datasets import load_dataset
# Thay token bằng Hugging Face Token của bạn
HF_TOKEN = "your_hf_token_here"… See the full description on the dataset page: https://huggingface.co/datasets/Khahn-nh/GTZAN-Dataset-Music-Genre-Classification.Color_Obj_Classification_20261003_164920This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_shoulder_pan.pos",
"left_shoulder_lift.pos",
"left_elbow_flex.pos",
"left_wrist_flex.pos",
"left_wrist_yaw.pos",
"left_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/joonyoung-seong/Color_Obj_Classification_20261003_164920.jigsaw-unintended-bias-in-toxicity-classificationColor_Obj_Classification_20261003_163617This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_shoulder_pan.pos",
"left_shoulder_lift.pos",
"left_elbow_flex.pos",
"left_wrist_flex.pos",
"left_wrist_yaw.pos",
"left_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/joonyoung-seong/Color_Obj_Classification_20261003_163617.Color_Obj_Classification_20261003_164359This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_shoulder_pan.pos",
"left_shoulder_lift.pos",
"left_elbow_flex.pos",
"left_wrist_flex.pos",
"left_wrist_yaw.pos",
"left_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/joonyoung-seong/Color_Obj_Classification_20261003_164359.car-classification-tabular-datasetcycic_classificationhttps://storage.googleapis.com/ai2-mosaic/public/cycic/CycIC-train-dev.zip
https://colab.research.google.com/drive/16nyxZPS7-ZDFwp7tn_q72Jxyv0dzK1MP?usp=sharing
@article{Kejriwal2020DoFC,
title={Do Fine-tuned Commonsense Language Models Really Generalize?},
author={Mayank Kejriwal and Ke Shen},
journal={ArXiv},
year={2020},
volume={abs/2011.09159}
}
added for
@article{sileo2023tasksource,
title={tasksource: Structured Dataset Preprocessing Annotations for Frictionless Extreme… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/cycic_classification.mqtl-classification-datasets-v1mqtl-classification-dataset-binned-200eval_smolvla_bi-so101-fruits-classificationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_so101_follower",
"total_episodes": 2,
"total_frames": 4987,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jy13/eval_smolvla_bi-so101-fruits-classification.WildChat-Legal-Classification
WildChat Legal Classification
English multi-turn conversations from WildChat-1M, silver-labeled with GPT-5.6 Sol at medium reasoning for legal-guidance detection, request typing, and legal topic assignment. This release uses the v3 taxonomy, which adds TAX_LAW as a standalone topic.
Four Hub configs are published on this repo:
Config
Rows
Description
legal
972
seeks_legal_guidance=true
nonlegal
1,150
seeks_legal_guidance=false
total
2,122
Full merged pool (former… See the full description on the dataset page: https://huggingface.co/datasets/AmirMohseni/WildChat-Legal-Classification.
