Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K199 likes11k downloads2y agoHugging Face02bitext /Bitext-retail-ecommerce-llm-chatbot-training-dataset Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.textquestion-answering10K<n<100K19 likes2.8k downloads2y agoHugging Face03bitext /Bitext-events-ticketing-llm-chatbot-training-dataset Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.textquestion-answering10K<n<100K1 likes2.2k downloads2y agoHugging Face04bitext /Bitext-telco-llm-chatbot-training-dataset Bitext - Telco Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [telco] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-telco-llm-chatbot-training-dataset.textquestion-answering10K<n<100K3 likes435 downloads2y agoHugging Face05bitext /Bitext-insurance-llm-chatbot-training-dataset Bitext - Insurance Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [insurance] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-insurance-llm-chatbot-training-dataset.textquestion-answering10K<n<100K9 likes337 downloads2y agoHugging Face06bitext /Bitext-travel-llm-chatbot-training-dataset Bitext - Travel Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Travel] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-travel-llm-chatbot-training-dataset.textquestion-answering10K<n<100K4 likes212 downloads2y agoHugging Face07SpX-DAC /training_datatextn<1K0 likes197 downloads10mo agoHugging Face08bitext /Bitext-hospitality-llm-chatbot-training-dataset Bitext - Hospitality Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [hospitality] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-hospitality-llm-chatbot-training-dataset.textquestion-answering10K<n<100K1 likes153 downloads2y agoHugging Face09bitext /Bitext-mortgage-loans-llm-chatbot-training-dataset Bitext - Mortgage and Loans Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Mortgage and Loans] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-mortgage-loans-llm-chatbot-training-dataset.textquestion-answering10K<n<100K5 likes144 downloads2y agoHugging Face10bitext /Bitext-wealth-management-llm-chatbot-training-dataset Bitext - Wealth Management Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Wealth Management] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-wealth-management-llm-chatbot-training-dataset.textquestion-answering10K<n<100K3 likes126 downloads2y agoHugging Face11bitext /Bitext-media-llm-chatbot-training-dataset Bitext - Media Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [media] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-media-llm-chatbot-training-dataset.textquestion-answering10K<n<100K0 likes117 downloads2y agoHugging Face12bitext /Bitext-restaurants-llm-chatbot-training-dataset Bitext - Restaurants Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [restaurants] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-restaurants-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes110 downloads2y agoHugging Face13bhumika-tewari-282006 /ayurvedic-affinity-training-data Ayurvedic Affinity Training Data The real, processed protein-ligand binding-affinity training set used to train the companion model bhumika-tewari-282006/ayurvedic-drug-discovery-affinity-model (RandomForest regressor for pKd prediction). What this is 181 real protein-ligand complexes from PDBBind v2013-core, each with: a real PDB structure identifier (pdb_code) the real ligand SMILES the real experimentally-measured binding affinity (pkd) 39 real, computed… See the full description on the dataset page: https://huggingface.co/datasets/bhumika-tewari-282006/ayurvedic-affinity-training-data.tabulartabular-regressionn<1K0 likes70 downloads8d agoHugging Face14violetxi /chess_puzzle_training_datasets_lt-2400 Chess puzzle training datasets: rating below 2400 This is a filtered derivative of pavelslab-nyu/chess_puzzle_training_datasets. Every retained row satisfies the exact condition: Rating < 2400 Rating is the Lichess puzzle rating, not the Elo of either player in the source game. The original column names, column order, directory layout, and CSV schemas are preserved. As in the upstream repository, Hugging Face discovers all three CSVs as one default configuration with one train… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/chess_puzzle_training_datasets_lt-2400.tabulartext-generation100K<n<1M0 likes53 downloads2mo agoHugging Face15wdouglass078 /Bitext-retail-ecommerce-llm-chatbot-training-dataset Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to… See the full description on the dataset page: https://huggingface.co/datasets/wdouglass078/Bitext-retail-ecommerce-llm-chatbot-training-dataset.textquestion-answering10K<n<100K0 likes44 downloads17d agoHugging Face16jason1966 /ahmedmohamed2003_cafe-sales-dirty-data-for-cleaning-training Cafe Sales - Dirty Data for Cleaning Training Dirty Cafe Sales Dataset Dataset Info Source: Kaggle Original Size: 0.11 MB Kaggle Downloads: 39,831 Files: 1 Files dirty_cafe_sales.csv Mirrored from Kaggle text10K<n<100K0 likes43 downloads6mo agoHugging Face17Geoweaver /ozone_training_data Ozone Training Data Dataset Summary The Ozone training dataset contains information about ozone levels, temperature, wind speed, pressure, and other related atmospheric variables across various geographic locations and time periods. It includes detailed daily observations from multiple data sources for comprehensive environmental and air quality analysis. Geographic coordinates (latitude and longitude) and timestamps (month, day, and hour) provide spatial and temporal… See the full description on the dataset page: https://huggingface.co/datasets/Geoweaver/ozone_training_data.tabular1M<n<10M2 likes41 downloads2y agoHugging Face18arunsrajan /godot4_training_data_sphinx_v2text10K<n<100K0 likes41 downloads6mo agoHugging Face19Dluvhugging /cc-tool-merchant-training-datatext10K<n<100K0 likes38 downloads24d agoHugging Face20ljoaql /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/ljoaql/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K0 likes34 downloads6mo agoHugging Face21Kubermatic /cncf-question-and-answer-dataset-for-llm-training CNCF QA Dataset for LLM Tuning Description This dataset, named cncf-qa-dataset-for-llm-tuning, is designed for fine-tuning large language models (LLMs) and is formatted in a question-answer (QA) style. The data is sourced from PDF and markdown (MD) files extracted from various project repositories within the CNCF (Cloud Native Computing Foundation) landscape. These files were processed and converted into a QA format to be fed into the LLM model. The dataset includes the… See the full description on the dataset page: https://huggingface.co/datasets/Kubermatic/cncf-question-and-answer-dataset-for-llm-training.text10K<n<100K3 likes33 downloads2y agoHugging Face22SID2702 /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/SID2702/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K0 likes32 downloads6mo agoHugging Face23UniDataPro /llm-training-dataset LLM Fine-Tuning Dataset - 4,000,000+ logs, 32 languages The dataset contains over 4 million+ logs written in 32 languages and is tailored for LLM training. It includes log and response pairs from 3 models, and is designed for language models and instruction fine-tuning to achieve improved performance in various NLP tasks - Get the data Models used for text generation: GPT-3.5 GPT-4 Uncensored GPT Version (is not included inthe sample) Languages in… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/llm-training-dataset.texttext-generation1K<n<10K3 likes31 downloads2mo agoHugging Face24ClarusC64 /clinical-quad-investigator-turnover-training-reset-protocol-deviations-data-lag-v0.1Clinical Quad Investigator Turnover Training Reset Protocol Deviations Data Lag v0.1 Each row is a site monthly snapshot. Core quad Investigator turnoverTraining resetProtocol deviationsData lag Target label_primary_fail_next_90d Files data/train.csvdata/tester.csvscorer.py Evaluation Run model on data/tester.csvReturn predictions row alignedScore with scorer.py License MIT This dataset identifies a measurable coupling pattern associated with systemic instability. The sample demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-investigator-turnover-training-reset-protocol-deviations-data-lag-v0.1.tabulartext-classificationn<1K0 likes31 downloads8mo agoHugging Face25mohamed0071 /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/mohamed0071/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K0 likes31 downloads4d agoHugging Face26harshgarg2006 /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/harshgarg2006/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K0 likes31 downloads3d agoHugging Face27abhi23457 /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/abhi23457/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K0 likes29 downloads1mo agoHugging Face28achrafgasmi /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/achrafgasmi/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K0 likes28 downloads8mo agoHugging Face29abullard1 /germeval-2025-harmful-content-detection-training-dataset GermEval 2025 Harmful Content Detection - Training Sets (Call to Action • Attacks on Democratic Basic Order • Violence) Author: Samuel Ruairí Bullard - University of Regensburg Models: Model Zoo (Gradio Space) Base model: LSX-UniWue/ModernGBERT_134M Competition: GermEval 2025 Shared Task Collection: GermEval 2025 Contribution CollectionabullardUR@GermEval Shared Task 2025 Submission Dataset Summary This repository republishes the training splits used… See the full description on the dataset page: https://huggingface.co/datasets/abullard1/germeval-2025-harmful-content-detection-training-dataset.texttext-classification10K<n<100K0 likes27 downloads1y agoHugging Face30leigangqu /DPT-T2I_training_datatext100K<n<1M1 likes26 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.