datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cbam-test-data
CBAM Test Data
Rows for CBAM reporting tests: declaration id, CN code, quantity, unit, embedded emissions, and country of origin. Synthetic values.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
cbam-small.json / cbam-small.csv — 100 rows (documents: 25)
cbam-medium.json / cbam-medium.csv — 2,000 rows (documents: 250)
cbam-large.json / cbam-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size always… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/cbam-test-data.2025_Virtual_Cell_Challenge_Test_Datatest-dataset-v1iata-awb-test-data
IATA Air Waybill Test Data
Air waybill numbers combining a 3-digit airline prefix, a 7-digit serial, and a MOD-7 check digit (IATA Resolution 600a).
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
iata-awb-small.json / iata-awb-small.csv — 100 rows (documents: 25)
iata-awb-medium.json / iata-awb-medium.csv — 2,000 rows (documents: 250)
iata-awb-large.json / iata-awb-large.csv — 20,000 rows (documents: 2,500)
Deterministic:… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/iata-awb-test-data.container-test-data
Container Number Test Data
Shipping container numbers with a correctly computed ISO 6346 check digit. Owner codes are synthetic.
Free to use under CC0-1.0 — public domain dedication, no attribution required.
Files
container-small.json / container-small.csv — 100 rows (documents: 25)
container-medium.json / container-medium.csv — 2,000 rows (documents: 250)
container-large.json / container-large.csv — 20,000 rows (documents: 2,500)
Deterministic: the same size… See the full description on the dataset page: https://huggingface.co/datasets/StanzaAPI/container-test-data.ml_data_test_detection_bank_transaction_frauds_unbalanced
ML Data Test Detection Bank Transaction Frauds Unbalanced
The project provides a quick and accessible dataset designed for learning and experimenting with machine learning algorithms, specifically in the context of detecting fraudulent bank transactions. It is intended for practicing and applying concepts such as Random Forest, Support Vector Machines (SVM), and Synthetic Minority Over-sampling Technique (SMOTE) to address unbalanced classification problems.
Note: This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/roberto-armas/ml_data_test_detection_bank_transaction_frauds_unbalanced.test-datasethf-free-maxxing-test-datahate_speech_open_data_original_class_test_setnsf-test-awards-dataNSF awards data for fiscal years 2024 and 2025.
Pulled from https://resources.research.gov/common/webapi/awardapisearch-v1.htm
very-test-dataset-2
My very good dataset
This dataset was carefully crafted in my home with a lot of coffee
By thomwolf
Test_Datasettest-datasettest_dataset2
Название датасета
Краткое описание датасета для предсказания сердечных заболеваний
Описание
Этот датасет содержит клинические параметры пациентов для диагностики сердечных заболеваний. Всего в наборе данных представлено 14 медицинских признаков и целевая переменная (диагноз).
Структура данных
Данные представлены в формате CSV со следующими столбцами:
Unnamed: 0 (int64) - технический индекс строки
Age (int64) - возраст пациента в годах
Sex (int64) - пол… See the full description on the dataset page: https://huggingface.co/datasets/Barzabel777/test_dataset2.very-test-dataset
My great dataset
test-dataset1TestDatatest-xml-data
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/c17hawke/test-xml-data.test_datasetbank-churn-train-test-datasetprompt_reverse_engineering_code_dataset_O3_arm_O3_advanced_custom_testtest-datasettest-volumes-dataset1_test_split_dataset
Mock Product Reviews Dataset
Dataset Description
A synthetic product review dataset for text classification and sentiment analysis tasks. The dataset contains user reviews across multiple product categories with ratings, sentiment labels, and metadata.
Dataset Summary
Total samples: 300
Train split: 210 samples (70.0%)
Validation split: 45 samples (15.0%)
Test split: 45 samples (15.0%)
Dataset Structure
Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/ldmLDM77/1_test_split_dataset.test-datatest-data2TestDataSub1-thinktest_data_v0.4Test_dataset_3
Hellow this is my dataset
columns:
emotion:
text:
test_vector_dataset
