datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
energy-consumption-hourly-spainenergy-consumption-weather-hourly-spainosworld_tasks_filesheloc
HELOC (Home Equity Line of Credit)
The HELOC dataset from FICO.
Each entry in the dataset is a line of credit, typically offered by a bank as a percentage of home equity (the difference between the current market value of a home and its purchase price).
The customers in this dataset have requested a credit line in the range of $5,000 - $150,000.
The fundamental task is to use the information about the applicant in their credit report to predict whether they will repay their HELOC… See the full description on the dataset page: https://huggingface.co/datasets/vitaliykinakh/heloc.bridgev2-vita-toykitchen-manifests
BridgeV2 VITA ToyKitchen-like Manifests
This repository contains manifest files for a reconstructed VITA-style BridgeV2 ToyKitchen-like pick-and-place subset.
Source dataset
The source dataset is:
Gaugou/BridgeV2
This repository does not duplicate the original BridgeV2 videos. It provides episode IDs and metadata for selecting the subset from the source dataset.
Split
Train: 2,986 episodes
Test: 287 episodes
Total selected: 3,273 episodes
Selection… See the full description on the dataset page: https://huggingface.co/datasets/praedico/bridgev2-vita-toykitchen-manifests.synthetic-fraud-detectionvithsd
Dataset Card for Dataset Name
ViTHSD: Vietnamese Targeted Hate Speech Detection
Dataset Details
A new version of Vietnamese Hate Speech Detection with target-oriented labels. Each comment can have multiple targets, each target has a hatred level indicating the hateful: CLEAN, OFFENSIVE, and HATE.
The targets are: Individuals, Groups, Religion/creed, Race/ethnicity, and Politics
Uses
Directly load and use the dataset from hugging face:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/sonlam1102/vithsd.csgosick
Thyroid Disease Dataset
Please refer to original source for more details.
Dataset Description
The Thyroid Disease dataset comprises medical records related to thyroid conditions. Supplied by the Garavan Institute and J. Ross Quinlan from the New South Wales Institute, Sydney, Australia in 1987, this dataset is widely used for diagnosing thyroid disorders. It contains 3,772 instances with 30 attributes, including both continuous and discrete features.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/vitaliykinakh/sick.gen_image_wordnet_preferencesThis dataset contains generated images. See the associated Hugging Face Collection for examples and additional details: Generated Image Wordnet
africa-synth-cancer-cancer-mortality-vital-registration-all
Cancer Mortality Vital Registration | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cancer-cancer-mortality-vital-registration-all.fmcw-vital-signsVITHSD
Dataset Card for VITHSD
1. Dataset Summary
VITHSD (Vietnamese Targeted Hate Speech Detection) contains 10,000 Vietnamese social‐media comments annotated for hate toward five target categories:
individual
groups
religion/creed
race/ethnicity
politics
Each target is labeled on a 3‐point scale (e.g., 0 = no hate, 1 = offensive, 2 = hateful). In this unified version, all splits are combined into one CSV with an extra type column indicating train / dev / test.
2.… See the full description on the dataset page: https://huggingface.co/datasets/visolex/VITHSD.travelSource
vi_term_definitionclinical_trial_patient_vitals
