Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pseudolab /huggingface-krew-hackathon2023textn<1K2 likes1.2k downloads3y agoHugging Face02traversaal-ai-hackathon /hotel_datasetsimage1K<n<10K3 likes203 downloads3y agoHugging Face03somosnlp-hackathon-2022 /Axolotl-Spanish-Nahuatl Axolotl-Spanish-Nahuatl : Parallel corpus for Spanish-Nahuatl machine translation Dataset Collection In order to get a good translator, we collected and cleaned two of the most complete Nahuatl-Spanish parallel corpora available. Those are Axolotl collected by an expert team at UNAM and Bible UEDIN Nahuatl Spanish crawled by Christos Christodoulopoulos and Mark Steedman from Bible Gateway site. After this, we ended with 12,207 samples from Axolotl due to misalignments and… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/Axolotl-Spanish-Nahuatl.tabulartranslation10K<n<100K16 likes98 downloads3y agoHugging Face04somosnlp-hackathon-2023 /suicide-comments-es Dataset Summary The dataset consists of comments on Reddit, Twitter, and inputs/outputs of the Alpaca dataset translated to Spanish language and classified as suicidal ideation/behavior and non-suicidal. Dataset Structure The dataset has 10050 rows (777 considered as Suicidal Ideation/Behavior and 9273 considered Not Suicidal). Dataset fields Text: User comment. Label: 1 if suicidal ideation/behavior; 0 if not suicidal comment. Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/suicide-comments-es.texttext-classification10K<n<100K5 likes59 downloads4y agoHugging Face05somosnlp-hackathon-2023 /Habilidades_Agente_v1 Description Español: Presentamos un conjunto de datos que presenta tres partes principales: 1. Dataset sobre habilidades blandas. 2. Dataset de conversaciones empresariales entre agentes y clientes. 3. Dataset curado de Alpaca en español: Este dataset toma como base el dataset https://huggingface.co/datasets/somosnlp/somos-alpaca-es, y fue curado con la herramienta Argilla, alcanzando 9400 registros curados. Los datos están estructurados en torno a un método que se describe… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/Habilidades_Agente_v1.texttext-generation10K<n<100K22 likes47 downloads3y agoHugging Face06somosnlp-hackathon-2025 /es-paremias-variantesParemias con sus variantes para entrenar modelo de embeddings. Obtenidos de web: https://cvc.cervantes.es/lengua/refranero/listado.aspx textsentence-similarityn<1K0 likes40 downloads1y agoHugging Face07somosnlp-hackathon-2022 /spanish-poetry-datasetThis dataset was previously created in Kaggle by Andrea Morales Garzón. Link Kaggle text1K<n<10K5 likes39 downloads5y agoHugging Face08somosnlp-hackathon-2025 /es-paremias-236236 paremias de Cultura Castellana. Tipos de paremia: Refrán,Refrán meteorológico, Proverbio, Frase proverbial, Aforismo, Locución proverbial, Dialogismo Obtenidos de web: https://cvc.cervantes.es/lengua/refranero/listado.aspx Distribución de tipos: textn<1K0 likes38 downloads1y agoHugging Face09wseo /huggingface-krew-hackathon23textn<1K0 likes34 downloads3y agoHugging Face10dfuc1302 /demo_hackathon Demo Hackathon Jailbreak Dataset Curated English dataset for calibrated jailbreak classification. text1K<n<10K0 likes32 downloads4d agoHugging Face11somosnlp-hackathon-2022 /comentarios_depresivosLa base de datos consta de una cantidad de 192 347 filas de datos para el entrenamiento, 33 944 para las pruebas y 22630 para la validación. Su contenido está compuesto por comentarios suicidas y comentarios normales de la red social Reddit traducidos al español, y obtenidos de la base de datos: Suicide and Depression Detection de Nikhileswar Komati, la cual se puede visualizar en la siguiente dirección: https://www.kaggle.com/datasets/nikhileswarkomati/suicide-watch Autores Danny Vásquez… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/comentarios_depresivos.text100K<n<1M4 likes31 downloads5y agoHugging Face12somosnlp-hackathon-2026 /Onexe-QA-Dataset Dataset Card: Canarian Linguistic Evaluation Dataset (QA without Answers) Dataset Summary This dataset has been designed specifically for evaluating the dialectal, linguistic, and cultural understanding of Large Language Models (LLMs) within the context of Canarian Spanish. It contains 4,683 evaluation questions based on the official lexicon of the Academy of Canarian Language (Academia Canaria de la Lengua - ACL). Each record presents a linguistic query phrased… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/Onexe-QA-Dataset.textquestion-answering1K<n<10K0 likes31 downloads4mo agoHugging Face13somosnlp-hackathon-2026 /somosnlp-2026-aerospace Dataset Card: Conjunto de Datos Aeroespacial y Cultural Completo Resumen del Dataset Este conjunto de datos ha sido diseñado específicamente para la evaluación cultural, lingüística y de alineación de Modelos de Lenguaje (LLMs) en el ámbito iberoamericano, con un foco especial en la historia aeroespacial, técnica, científica e histórica. Contiene 1.716 interacciones de tipo conversacional (multi-turn) distribuidas en múltiples países de habla hispana y portuguesa.… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/somosnlp-2026-aerospace.texttext-generation1K<n<10K0 likes30 downloads4mo agoHugging Face14somosnlp-hackathon-2022 /poems-esDataset descargado de la página kaggle.com. El archivo original contenía información en inglés y posteriormente fue traducida para su uso. El dataset contiene las columnas: Autor: corresponde al autor del poema. Contenido: contiene todo el poema. Nombre del poema: contiene el título del poema. Años: corresponde al tiempo en que fue hecho el poema. Tipo: contiene el tipo que pertenece el poema. textn<1K3 likes29 downloads5y agoHugging Face15somosnlp-hackathon-2025 /cenia-team-sabiduriapopulartextquestion-answeringn<1K0 likes28 downloads1y agoHugging Face16somosnlp-hackathon-2022 /scientific_papers_entext1K<n<10K1 likes27 downloads5y agoHugging Face17somosnlp-hackathon-2025 /es-refranes-145145 refranes cortos de Cultura Castellana. Obtenidos de web: https://psicologiaymente.com/reflexiones/refranes-cortos Xavier Molina. (2019, octubre 4). 145 refranes cortos muy populares (y su significado). Portal Psicología y Mente. https://psicologiaymente.com/reflexiones/refranes-cortos textn<1K0 likes23 downloads1y agoHugging Face18somosnlp-hackathon-2025 /es-refranes-103103 refranes cortos de Cultura Castellana. Obtenidos de web: https://www.curiosidario.es/refranes-populares-espanoles/ textn<1K0 likes23 downloads1y agoHugging Face19somosnlp-hackathon-2025 /es-refranes-datasettexttext-generationn<1K0 likes21 downloads1y agoHugging Face20somosnlp-hackathon-2022 /Dataset-Acoso-Twitter-Es UNL: Universidad Nacional de Loja Miembros del equipo: Anderson Quizhpe Luis Negrón David Pacheco Bryan Requenes Paul Pasaca textn<1K2 likes18 downloads5y agoHugging Face21Saumya536 /Hackathon_Datasettext1K<n<10K0 likes18 downloads2y agoHugging Face22arychaud /piimask-hackathon mysqlclient This project is a fork of MySQLdb1. This project adds Python 3 support and fixed many bugs. PyPI: https://pypi.org/project/mysqlclient/ GitHub: https://github.com/PyMySQL/mysqlclient Support Do Not use Github Issue Tracker to ask help. OSS Maintainer is not free tech support When your question looks relating to Python rather than MySQL: Python mailing list python-list Slack pythondev.slack.com Or when you have question about MySQL: MySQL Community on… See the full description on the dataset page: https://huggingface.co/datasets/arychaud/piimask-hackathon.text10K<n<100K0 likes16 downloads3y agoHugging Face23somosnlp-hackathon-2022 /scientific_papers_en_estext1K<n<10K1 likes15 downloads5y agoHugging Face24awhedon /hackathon-instruction-finetuningtextn<1K0 likes15 downloads3y agoHugging Face25somosnlp-hackathon-2025 /es-paremias-variantes-antonimosAmpliación del dataset es-paremias-variantes con la columna "Frases Antonimas" que pretende ser una frase que tiene el significado completamente opuesto al de la frase variante. ❗Esta columna se ha generado de forma automática utilizando el modelo Phi-4 a través de LMStudio. La inspección manual de algunos ejemplos valida que sea una frase que mantiene el significado opuesto, pero no todos los ejemplos han sido revisados. Sientete libre de abrir pull-request con mejoras a este dataset ✨textsentence-similarityn<1K0 likes15 downloads1y agoHugging Face26AFZAL0008 /hackathontext10K<n<100K0 likes14 downloads2y agoHugging Face27KuuwangE /hack-a-thon-resulttabularn<1K0 likes14 downloads1y agoHugging Face28somosnlp-hackathon-2022 /disco_spanish_poetry DISCO: Diachronic Spanish Sonnet Corpus The Diachronic Spanish Sonnet Corpus (DISCO) contains sonnets in Spanish in CSV, between the 15th and the 20th centuries (4303 sonnets by 1215 authors from 22 different countries). It includes well-known authors, but also less canonized ones. This is a CSV compilation taken from the plain text corpus v4 published on git https://github.com/pruizf/disco/tree/v4. It includes the title, author, age and text metadata. text1K<n<10K7 likes13 downloads5y agoHugging Face29somosnlp-hackathon-2022 /scientific_papers_estext1K<n<10K0 likes12 downloads5y agoHugging Face30Lancelot53 /hackathontext10K<n<100K0 likes12 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.