datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kicky-ai-spf
FUT-HEROS SPF — COCO instance-segmentation dataset (ball / player / goal)
Auto-labelled football frames (COCO format) for ball/player/goal instance segmentation.
Labels generated by SAM3 + NVIDIA LocateAnything-3B (zero manual annotation), from a
single fixed-camera amateur session. Clip-level, goal-stratified train/valid/test split.
train/, valid/, test/ each have _annotations.coco.json + frames.
Classes: 1 ball, 2 player, 3 goal. Part of the FUT-HEROS project.
informes_discriminacion_gitana
Resumen del dataset
Se trata de un dataset en español, extraído del centro de documentación de la Fundación Secretariado Gitano, en el que se presentan distintas situaciones discriminatorias acontecidas por el pueblo gitano. Puesto que el objetivo del modelo es crear un sistema de generación de actuaciones que permita minimizar el impacto de una situación discriminatoria, se hizo un scrappeo y se extrajeron todos los PDFs que contuvieron casos de discriminación con el formato… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/informes_discriminacion_gitana.hotel_datasetsux-crime-scene-traces
🔎 UX Crime Scene — Investigation Traces
Real agent traces from UX Crime Scene,
a film-noir detective that investigates UI screenshots as crime scenes — built for the
Build Small Hackathon (Gradio × Hugging Face).
Each row is one real investigation: the input screenshot the user dropped, and the
raw structured verdict Qwen2.5-VL-7B returned — the crimes it found, the bounding
box of each guilty element, the testimony, the severity, and the final grade.
The set spans different… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/ux-crime-scene-traces.memes_instagram_chilenos_es_small
memes_instagram_chilenos_es_small
A dataset designed to train and evaluate vision-language models on Chilean meme understanding, with a strong focus on cultural context and local humor, built for the Somos NLP Hackathon 2025.
Introduction
Memes are rich cultural artifacts that encapsulate humor, identity, and social commentary in visual formats. Yet, most existing datasets focus on English-language content or generic humor detection, leaving culturally grounded… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/memes_instagram_chilenos_es_small.250-ez_2_ai-so100_hack_medexcavator-hackathongradio-agents-mcp-hackathon-certificates
Dataset Card for "gradio-agents-mcp-hackathon-certificates"
More Information needed
neural-earth-fields-hackathon
Neural Earth Fields Hackathon
Land cover
CGLC-MODIS-LCZ-100m.cog.tif is a Cloud Optimized GeoTIFF version of A hybrid 100-m global land cover dataset with Local Climate Zones for WRF by Matthias Demuzere, Cenlin He, Alberto Martilli, and Andrea Zonato (2023).
The source dataset is licensed under CC BY 4.0.
The file contains one uint8 class ID per pixel on the source's nominal 100 m EPSG:4326 grid.
It uses ZSTD compression, 512 × 512 tiles, and mode-resampled… See the full description on the dataset page: https://huggingface.co/datasets/taylor-geospatial/neural-earth-fields-hackathon.moku-blog-assetswatercolour-life-klein-4b-lora-80tiny-dispatch-coach-media
Tiny Dispatch Coach Demo Media
Public demo assets for the Build Small Hackathon submission:
Space: https://huggingface.co/spaces/build-small-hackathon/tiny-dispatch-coach
Runtime: https://build-small-hackathon-tiny-dispatch-coach.hf.space/
Trace dataset: https://huggingface.co/datasets/build-small-hackathon/tiny-dispatch-coach-traces
Real demo recording: tiny-dispatch-coach-real-demo.webm
The demo recording is a real browser capture of the public Space: it opens the
runtime… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/tiny-dispatch-coach-media.kirana-invoice-train-data
Kirana Invoice Training Data — Indian FMCG
Training dataset for the Kirana Detective project — an AI pipeline that audits distributor invoices for Indian kirana (grocery) stores. The repository contains two distinct sub-datasets used to fine-tune two separate models.
Dataset Summary
Sub-dataset
Purpose
Size
Format
synthetic_invoices/
OCR fine-tuning (MiniCPM-V)
500 images + annotations
PNG + JSONL
fmcg_catalog.json
Product name normalization… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/kirana-invoice-train-data.ibero-characters-es
Conjunto de datos de personajes de mitos y leyendas iberoamericanos.
⚠️ Este dataset se encuentra en desarrollo activo. Se planea expandir significativamente el número de registros y mejorar la cobertura de imágenes.
📚 Descripción
Dataset de personajes míticos y legendarios de Iberoamérica, diseñado para preservar y promover el patrimonio cultural a través de la inteligencia artificial.
🌟 Motivación e Impacto
📱 Preservación Digital: Conservación del… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/ibero-characters-es.202-cokkiri-robo-barista-vla-datasetmcp-birthday-hackathon-certificatesTensor-Network-Hackathon-2024fatyoshi-dreambooth-hackathon-imagesdreambooth-hackathon-nala
Dataset Card for "dreambooth-hackathon-nala"
More Information needed
patrimonio-gastronomico-hispano
Patrimonio Gastronómico Hispano
Patrimonio Gastronómico Hispano
Descripción
Este dataset contiene una colección estructurada de recetas tradicionales ecuatorianas en español. Incluye recetas detalladas con ingredientes, instrucciones paso a paso, información nutricional, y enlaces a videos demostrativos, preservando el rico patrimonio culinario de Ecuador.
Propósito
El dataset fue creado para:
Preservar y digitalizar el conocimiento… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/patrimonio-gastronomico-hispano.dreambooth-hackathon-images-srkman
Dataset Card for "dreambooth-hackathon-images-srkman"
More Information needed
dreambooth-hackathon-images-srkman-2
Dataset Card for "dreambooth-hackathon-images-srkman-2"
More Information needed
P4Ms-hackathon-vision-taskdreambooth-hackathon-images
Dataset Card for "dreambooth-hackathon-images"
More Information needed
ordfts-hackathon-pneuma-vehicles-segmentation
ORD for the Sciences Hackathon - Vehicles Detection
[!CAUTION]
This project is an example of a hackathon project. The quality of the data produced has not been evaluated. Its goal is to provide an example on how a dataset can be update to Hugginface.
This is an example of a hackathon project presented to ORD for the sciences hackathon using the openly available pNeuma vision dataset.
Go here if you wanna know more about the hackathon
EPFL pNeuma project… See the full description on the dataset page: https://huggingface.co/datasets/katospiegel/ordfts-hackathon-pneuma-vehicles-segmentation.RoboDomi-RobotLoadDishwasher
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/RoboDomi-RobotLoadDishwasher.dreambooth-hackathon-rick-and-morty-images
Dataset Card for "dreambooth-hackathon-rick-and-morty-images"
More Information needed
dreambooth-hackathon-images
Dataset Card for "dreambooth-hackathon-images"
More Information needed
CXR_BioXAi_Hackathon_2024hackathon_png
