Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AiMijie /EC-Guide This repo is only used for dataset viewer. Please download from here. Amazon KDDCup 2024 Team ZJU-AI4H’s Solution and Dataset (Track 2 Top 2; Track 5 Top 5) The Amazon KDD Cup’24 competition presents a unique challenge by focusing on the application of LLMs in E-commerce across multiple tasks. Our solution for addressing Tracks 2 and 5 involves a comprehensive pipeline encompassing dataset construction, instruction tuning, post-training quantization, and inference… See the full description on the dataset page: https://huggingface.co/datasets/AiMijie/EC-Guide.textquestion-answering10K<n<100K2 likes208 downloads2y agoHugging Face02projecte-aina /GuiaCat Dataset Card for GuiaCat Dataset Summary GuiaCat is a dataset consisting of 5.750 restaurant reviews in Catalan, with 5 associated scores and a label of sentiment. The data was provided by GuiaCat and curated by the BSC. This work is licensed under a Creative Commons Attribution Non-commercial No-Derivatives 4.0 International License. Supported Tasks and Leaderboards This corpus is mainly intended for sentiment analysis. Languages The… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/GuiaCat.tabulartext-classification1K<n<10K0 likes193 downloads3y agoHugging Face03jojo-ai-mst /Myanmar-Tuberculosis-Guidelines-Instructions Myanmar Tuberculosis Guidelines Instructions A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages. Authors: Min Si Thu, Khin Myat Noe Abstract Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.imagequestion-answering1K<n<10K1 likes174 downloads5mo agoHugging Face04SmartQHSE /hazop-guidewords-reference-2026 Canonical landing page: https://www.smartqhse.com/datasets/hazop-guidewords-reference-2026 HAZOP Guidewords Reference 2026 — IEC 61882 with Node Examples HAZOP guidewords reference following IEC 61882:2016 and CCPS guidance. 7 standard guidewords (NO/NOT, MORE, LESS, AS WELL AS, PART OF, REVERSE, OTHER THAN) applied to key process parameters (flow, pressure, temperature, level, composition, reaction, rotation, operating mode) with ~20 worked examples each providing: deviation… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hazop-guidewords-reference-2026.textn<1K0 likes139 downloads4mo agoHugging Face05SnehaPriyaaMP /WCAG_Combined_guideline_datatext10K<n<100K0 likes98 downloads2y agoHugging Face06Guizhen /Puzzles_10ktext10K<n<100K0 likes95 downloads2y agoHugging Face07Pradeep016 /career-guidance-qa-dataset Dataset Card for Career Guidance Dataset Dataset Overview This dataset provides career guidance information for a variety of career roles. It includes questions and answers related to career roles such as "Data Scientist," "Software Engineer," "Product Manager," and many more. The dataset covers aspects like job responsibilities, required skills, career progression, salary expectations, and work environment. It is intended for use in building chatbot applications for… See the full description on the dataset page: https://huggingface.co/datasets/Pradeep016/career-guidance-qa-dataset.textquestion-answering1K<n<10K7 likes89 downloads2y agoHugging Face08GuillermoLlopis /ai-inference-emission-factors SOMA AI-Inference Emission and Resource Factors Ready-to-use carbon and water emission factors for estimating the footprint of AI inference (LLM API calls) in corporate sustainability inventories — built for CSRD / GHG Protocol Scope 3 Category 1 reporting. Every factor is derived from primary, cited sources (GPU energy benchmarks, grid carbon intensity registries, datacenter water-use studies); derivations are documented column-by-column below and in full in Supplementary S1 of… See the full description on the dataset page: https://huggingface.co/datasets/GuillermoLlopis/ai-inference-emission-factors.tabularn<1K0 likes82 downloads20d agoHugging Face09TylerHilbert /3D-Printable-Guitar-Modelstextn<1K0 likes54 downloads5mo agoHugging Face10guillermoruiz /RegTweetsThis is a dataset created from Tweets written in Spanish, particularly from México. The classification task is to chose the Mexican State where the message was written from. The label correspond to the state given the following table: State name Label State name Label Aguascalientes Aguascalientes Mexico Mexico Baja California BC Nayarit Nayarit Baja California Sur BCS Nuevo León NL Campeche Campeche Oaxaca Oaxaca Chiapas Chiapas Puebla Puebla Chihuahua Chihuahua Querétaro… See the full description on the dataset page: https://huggingface.co/datasets/guillermoruiz/RegTweets.texttext-classification100K<n<1M0 likes50 downloads2y agoHugging Face11guillermoruiz /MexEmojisThis is a dataset of tweets written in Mexican Spanish and the labels is an emoji describing the emotion. The region columns is the token for the Mexican state from where the message was written. The state token is one of the following: State name Label State name Label Aguascalientes Aguascalientes Mexico Mexico Baja California BC Nayarit Nayarit Baja California Sur BCS Nuevo León NL Campeche Campeche Oaxaca Oaxaca Chiapas Chiapas Puebla Puebla Chihuahua Chihuahua… See the full description on the dataset page: https://huggingface.co/datasets/guillermoruiz/MexEmojis.texttext-classification100K<n<1M0 likes48 downloads2y agoHugging Face12guidobenb /vcdbtexttoken-classification1K<n<10K0 likes45 downloads3y agoHugging Face13wolferussell14 /massive-guitar-8b0d4e massive-guitar-8b0d4e Synthetic sensors test data: 34 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/wolferussell14/massive-guitar-8b0d4e.tabularn<1K0 likes43 downloads25d agoHugging Face14hadro /green-books-travel-guides African American Travel Guides: The Green Book & Companion Directories (1930–1966) A unified, structured dataset of 113,827 business and lodging listings transcribed from 50 volumes of mid-20th-century African American travel guides, spanning 1930–1966. During the Jim Crow era, these guides told Black travelers which hotels, restaurants, tourist homes, service stations, and other businesses would serve them safely. This dataset brings The Negro Motorist Green Book together with… See the full description on the dataset page: https://huggingface.co/datasets/hadro/green-books-travel-guides.texttext-classification100K<n<1M1 likes41 downloads2mo agoHugging Face15ttn0011 /pageguide_guide_data PageGuide Dataset This repository contains the dataset for PageGuide, a browser extension that assists users in navigating webpages and locating information by grounding LLM answers directly in the HTML DOM. Project Page: pageguide.github.io Paper: PageGuide: Browser extension to assist users in navigating a webpage and locating information Code: github.com/tin-xai/pageguide Dataset Description The PageGuide evaluation utilizes several distinct datasets… See the full description on the dataset page: https://huggingface.co/datasets/ttn0011/pageguide_guide_data.tabularothern<1K0 likes35 downloads3mo agoHugging Face16Arrakark /scalared_guidelines Scalared Guidelines Description The following dataset is a collection of 7226 rules and guidelines related to digital design and digital verification. They can be used to create an "AI linter" for digital designs and testbenches. There are 10 fields in each row: title: a very short human-readable description of the guideline. Useful mostly for uniquifying the guidelines. statement: a single sentence of what a user should do to follow the guideline description: a… See the full description on the dataset page: https://huggingface.co/datasets/Arrakark/scalared_guidelines.text1K<n<10K1 likes34 downloads10mo agoHugging Face17GuiGel /meddocan Dataset Card for "meddocan" Dataset Summary A personal upload of the SPACC_MEDDOCAN corpus. The tokenization is made with the help of a custom spaCy pipeline. Supported Tasks and Leaderboards Name Entity Recognition Languages More Information Needed Dataset Structure Data Instances More Information Needed Data Fields The data fields are the same among all splits. Data Splits name train validation test… See the full description on the dataset page: https://huggingface.co/datasets/GuiGel/meddocan.texttoken-classification10K<n<100K0 likes32 downloads4y agoHugging Face18ClarusC64 /clinical-tpib-invariant-guided-next-intervention-prediction-v0.1What this dataset tests Given a patient’s manifold typepredict the top 3 next interventions that are most coherent. It rewards manifold-consistent moves constraint-aware choices cross-domain suggestions when warranted It penalizes repeating tolerance loops repeating paradoxical worseners ignoring contraindications choosing common care without manifold fit Labels coherent_top3 partially_coherent_top3 incoherent_top3 Suggested prompt wrapper System You propose the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-tpib-invariant-guided-next-intervention-prediction-v0.1.texttext-classificationn<1K0 likes31 downloads8mo agoHugging Face19lei-HuggingFace /guide-to-level-measurement-2021-edition-en-77708_combined_synthetic_datatext1K<n<10K0 likes27 downloads2y agoHugging Face20GuillermoTBB /gp-long-paragraphstabular10M<n<100M1 likes27 downloads2y agoHugging Face21guillemram97 /PEEP PEEP: Prompts, Extracted Entities with Privacy Paper: Controlling What You Share: Assessing Language Model Adherence to Privacy Preferences Dataset Summary PEEP is a multilingual dataset of 15,282 real user queries from the Wildchat dataset, annotated with extracted personal information and associated with synthetic privacy profiles. It is designed to support research on privacy-preserving language models, enabling controlled evaluation of models’ adherence to… See the full description on the dataset page: https://huggingface.co/datasets/guillemram97/PEEP.textquestion-answering10K<n<100K1 likes24 downloads1y agoHugging Face22electricsheepafrica /air-pollution-population-exposed-to-levels-exceeding-who-guideline-percentage-of-total-africa Air Pollution Population Exposed to Levels Exceeding WHO Guideline Percentage of Total Africa | Africa (World Health Organization) Size category: n<1K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/air-pollution-population-exposed-to-levels-exceeding-who-guideline-percentage-of-total-africa.tabulartabular-classificationn<1K0 likes24 downloads2mo agoHugging Face23nlp-guild /intent-recognition-biomedicalsource text1K<n<10K4 likes23 downloads4y agoHugging Face24GuillermoTBB /charles-dickens-text-classification Dataset Description This Dataset is designed to classify paragraphs of text as either written by Charles Dickens or generated to imitate various distinct writing styles. The primary use case is in the domain of literary analysis and text generation, where it can help distinguish between authentic Dickensian text and stylistic imitations. The dataset was created from paragraphs extracted from “Great Expectations” by Charles Dickens. The original text was taken from the Gutenberg… See the full description on the dataset page: https://huggingface.co/datasets/GuillermoTBB/charles-dickens-text-classification.texttext-classification1K<n<10K1 likes23 downloads2y agoHugging Face25Gopher-Lab /Tiktok_Chatgpt_Prompt_Guide 📲 Example Dataset: TikTok Scraper Tool 👉 Start Scraping TikTok: TikTok Scraper Tool ✨ Key Features ⚡ Instant Transcription – Turn any TikTok video into an AI-ready transcript 🎯 Metadata – Get the title, language, description, and video hashtags 🔗 URL-Based Access – Just drop in a TikTok video URL to start scraping 🧩 LLM-Ready Output – Receive clean JSON ready for agents, RAG, or AI tools 💸 Free Tier – Use up to 100 queries during the beta period 💫 Easy… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/Tiktok_Chatgpt_Prompt_Guide.textfeature-extraction1K<n<10K0 likes23 downloads1y agoHugging Face26keygate-ai /ai-model-evaluation-guide AI 模型选型与测评维度词典 版本:1.0.0|更新日期:2026-07-29 Keygate 是覆盖全球主流与前沿 AI 模型的测评、排行榜与选型平台。这份中英双语词典将语言、图像、视频与语音模型比较中常见的 18 项指标整理为结构化字段,帮助读者正确理解榜单、建立选型表,并减少不同资料之间的术语混用。 A bilingual data dictionary of 18 dimensions for evaluating and selecting leading language, image, video and speech models. 配套资料 Keygate 实时排行榜、模型详情与并排对比 GitHub:AI 模型测评与选型维度指南 公开评测基准索引 可下载的评测基准 CSV 数据内容 统一中英文指标名称,减少同一概念被不同译法混用。 明确数值应当“越高越好”还是“越低越好”。 区分输出速度与首段响应时间,避免把两个概念当成同一项。… See the full description on the dataset page: https://huggingface.co/datasets/keygate-ai/ai-model-evaluation-guide.textn<1K0 likes23 downloads2mo agoHugging Face27VamshiSurya /Career-Guidance 📚 Career Compass Instruction Dataset The Career Compass Instruction Dataset is a curated set of instruction-response pairs designed to train or evaluate conversational AI systems that provide personalized career guidance. It reflects real queries students may ask and informative answers from the Career Compass AI. ✨ Dataset Summary This dataset contains high-quality examples of how a conversational AI system can provide career-related guidance. Each row contains:… See the full description on the dataset page: https://huggingface.co/datasets/VamshiSurya/Career-Guidance.text1K<n<10K2 likes18 downloads1y agoHugging Face28ClarusC64 /clinical-guideline-strength-correspondence-v0.1 What this dataset tests Guideline strength must track evidence strength. Authority must not exceed data. Why it exists Guidelines often harden too early. Language outruns certainty. This set checks whether recommendation force matches evidence quality. Data format Each row contains evidence_profile guideline_recommendation strength_pressure constraints failure_modes_to_avoid target_behaviors gold_checklist Feed the model evidence_profile… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-guideline-strength-correspondence-v0.1.texttext-classificationn<1K0 likes16 downloads8mo agoHugging Face29GuillermoBarrio /dataset-electoral Electoral & Socioeconomic Dataset (Spain 2011-2019) - RAG Optimized Description / Descripción EN: This dataset contains unified electoral results and socioeconomic indicators (average income) at the census section level for the Spanish General Elections held between 2011 and 2019. It has been specifically processed and formatted with metadata strings to power Retrieval-Augmented Generation (RAG) systems. ES: Este dataset contiene los resultados electorales unificados e… See the full description on the dataset page: https://huggingface.co/datasets/GuillermoBarrio/dataset-electoral.tabular100K<n<1M0 likes16 downloads8mo agoHugging Face30Guillem21 /QA_electronics_datasettext100K<n<1M0 likes14 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.