Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K199 likes11k downloads2y agoHugging Face02openadmet /cyp-challenge-train-test CYP Challenge Train/Test Dataset A high-quality experimental dataset for predicting inhibition of the major drug-metabolizing Cytochrome P450 enzymes (CYP1A2, CYP2C9, CYP2D6, CYP3A4), released as part of the OpenADMET CYP Inhibition Blind Challenge. Blog post: Announcing OpenADMET’s CYP inhibition blind challenge Challenge Space: OpenADMET CYP Inhibition Blind Challenge Challenge period: August 17, 2026 - November 3, 2026 Produced by: OpenADMET CHANGELOG Updated… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/cyp-challenge-train-test.tabulartabular-regression10K<n<100K12 likes3.2k downloads17d agoHugging Face03bitext /Bitext-retail-ecommerce-llm-chatbot-training-dataset Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.textquestion-answering10K<n<100K19 likes2.8k downloads2y agoHugging Face04bitext /Bitext-events-ticketing-llm-chatbot-training-dataset Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.textquestion-answering10K<n<100K1 likes2.2k downloads2y agoHugging Face05UW-Madison-Lee-Lab /MMLU-Pro-CoT-Train-Labeled Dataset Details Modality: Text Format: CSV Size: 10K - 100K rows Total Rows: 84,098 License: MIT Libraries Supported: datasets, pandas, croissant Structure Each row in the dataset includes: question: The query posed in the dataset. answer: The correct response. category: The domain of the question (e.g., math, science). src: The source of the question. id: A unique identifier for each entry. chain_of_thoughts: Step-by-step reasoning steps leading to the answer. labels:… See the full description on the dataset page: https://huggingface.co/datasets/UW-Madison-Lee-Lab/MMLU-Pro-CoT-Train-Labeled.text10K<n<100K6 likes1.6k downloads2y agoHugging Face06openadmet /pxr-challenge-train-test PXR Challenge Train/Test Dataset A high-quality experimental dataset for predicting human Pregnane-X Receptor (PXR) induction, comprising over 11,000 compounds screened using a high-fidelity in-house assay. This is the largest publicly available PXR activity dataset, released as part of the OpenADMET PXR Induction Blind Challenge. Blog post: Announcing the Next OpenADMET Blind Challenge: Predicting PXR Induction Challenge Space: openadmet/pxr-challenge Challenge period: April 1… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/pxr-challenge-train-test.tabulartabular-regression10K<n<100K17 likes945 downloads12d agoHugging Face07openadmet /openadmet-expansionrx-challenge-train-data OpenADMET-ExpansionRx Challenge training dataset This dataset contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases. While optimising candidate molecules for their preclinical programs Expansion collected a variety of ADMET data for off-targets and properties of interest in the traditional game of “whack-a-mole” familiar to all drug hunters. Now, they’ve made the bold and generous decision to… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-train-data.tabular10K<n<100K10 likes536 downloads10mo agoHugging Face08bitext /Bitext-telco-llm-chatbot-training-dataset Bitext - Telco Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [telco] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-telco-llm-chatbot-training-dataset.textquestion-answering10K<n<100K3 likes435 downloads2y agoHugging Face09SmellsLikeAISpirit /plant-disease-trainimage10K<n<100K0 likes429 downloads7mo agoHugging Face10bitext /Bitext-insurance-llm-chatbot-training-dataset Bitext - Insurance Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [insurance] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-insurance-llm-chatbot-training-dataset.textquestion-answering10K<n<100K9 likes337 downloads2y agoHugging Face11jliang097 /FARM_training_test FARM Aerial Radio Map (ARM) Dataset Paper: FARM: Foundational Aerial Radio Map for Intelligent Low-Altitude Networking (https://arxiv.org/abs/2604.17362) Overview This repository releases the constructed ARM datasets based on ARM-Omni for FARM training, in-domain evaluation (D1-D10), and zero-shot evaluation (P1, F1, and A1). The dataset coverage is summarized below: Dataset Frequencies (GHz) Max Rx Height (m) Beamwidths Map Grid Size Volume D1 2.1… See the full description on the dataset page: https://huggingface.co/datasets/jliang097/FARM_training_test.tabularimage-to-image10K<n<100K1 likes329 downloads5mo agoHugging Face12MUG-V /MUG-V-Training-Samples MUG-V Training Samples Sample training dataset for the MUG-V 10B video generation model training framework. Dataset Description This dataset contains pre-processed training samples for quick-start validation and testing of the MUG-V Megatron-LM training pipeline. It includes: VideoVAE-encoded latents (8×8×8 compressed video representations) T5-XXL text features (4096-dim embeddings) Training metadata CSV (sample mapping and configuration) ⚠️ Note: This is a sample… See the full description on the dataset page: https://huggingface.co/datasets/MUG-V/MUG-V-Training-Samples.texttext-to-video1K<n<10K0 likes322 downloads1y agoHugging Face13previtus /STARCOP_allbands_Train1gated STARCOP dataset STARCOP dataset: Semantic Segmentation of Methane Plumes with Hyperspectral Machine Learning Models 🌈🛰️Authors: Vít Růžička, Gonzalo Mateo-Garcia, Luis Gómez-Chova, Anna Vaughan, Luis Guanter and Andrew Markham Fast data preview in: dataset_exploration.ipynb Main repository: github/spaceml-org/STARCOP Task: Methane is the second most important greenhouse gas contributor to climate change; at the same time its reduction has been denoted as one of the… See the full description on the dataset page: https://huggingface.co/datasets/previtus/STARCOP_allbands_Train1.imageimage-segmentation1K<n<10K3 likes302 downloads2y agoHugging Face14bitext /Bitext-travel-llm-chatbot-training-dataset Bitext - Travel Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Travel] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-travel-llm-chatbot-training-dataset.textquestion-answering10K<n<100K4 likes212 downloads2y agoHugging Face15SpX-DAC /training_datatextn<1K0 likes197 downloads10mo agoHugging Face16redmadrobot-rnd /pii_train Russian PII NER Training Dataset Dataset Description This is the training corpus for PII (Personally Identifiable Information) detection and Named Entity Recognition (NER) on Russian-language text. It targets guardrail and anonymization pipelines that must reliably find personal data (names, addresses, contacts) and Russian identity-document numbers (passport, SNILS, INN, OMS, etc.) in text. The corpus combines real, manually annotated examples from production… See the full description on the dataset page: https://huggingface.co/datasets/redmadrobot-rnd/pii_train.texttoken-classification10K<n<100K1 likes174 downloads1mo agoHugging Face17PersonaBias /Reverse-hybrid-correct-train-no-persona-meantabulartext-classification100K<n<1M0 likes157 downloads2mo agoHugging Face18PersonaBias /Reverse-hybrid-train-no-persona-meantabulartext-classification100K<n<1M0 likes157 downloads2mo agoHugging Face19bitext /Bitext-hospitality-llm-chatbot-training-dataset Bitext - Hospitality Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [hospitality] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-hospitality-llm-chatbot-training-dataset.textquestion-answering10K<n<100K1 likes153 downloads2y agoHugging Face20CaiYuanhao /OmniVCus-Train [NeurIPS 2025] OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions Dataset Description This dataset supports multi-modal constrol video generation. It contains ~80K data samples processed from 140K videos by our VideoCus-Factory pipeline. Each data sample includes the original video, text prompts, subject reference image, depth video, mask video, and motion video conditions. Here is a data example: Generated… See the full description on the dataset page: https://huggingface.co/datasets/CaiYuanhao/OmniVCus-Train.text100K<n<1M2 likes152 downloads9mo agoHugging Face21bitext /Bitext-mortgage-loans-llm-chatbot-training-dataset Bitext - Mortgage and Loans Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Mortgage and Loans] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-mortgage-loans-llm-chatbot-training-dataset.textquestion-answering10K<n<100K5 likes144 downloads2y agoHugging Face22harryxi /HelpSteer3-general-code-Shift-Qwen-2.5-1.5B-Instruct-dataset-train-generationstext1M<n<10M0 likes129 downloads1y agoHugging Face23bitext /Bitext-wealth-management-llm-chatbot-training-dataset Bitext - Wealth Management Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Wealth Management] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-wealth-management-llm-chatbot-training-dataset.textquestion-answering10K<n<100K3 likes126 downloads2y agoHugging Face24sebastian-hofstaetter /tripclick-training TripClick Baselines with Improved Training Data Establishing Strong Baselines for TripClick Health Retrieval Sebastian Hofstätter, Sophia Althammer, Mete Sertkan and Allan Hanbury https://arxiv.org/abs/2201.00365 tl;dr We create strong re-ranking and dense retrieval baselines (BERTCAT, BERTDOT, ColBERT, and TK) for TripClick (health ad-hoc retrieval). We improve the – originally too noisy – training data with a simple negative sampling policy. We achieve large gains over BM25 in the… See the full description on the dataset page: https://huggingface.co/datasets/sebastian-hofstaetter/tripclick-training.tabulartext-retrieval1M<n<10M1 likes125 downloads4y agoHugging Face25gsingh1-py /traintext1K<n<10K9 likes120 downloads2y agoHugging Face26bitext /Bitext-media-llm-chatbot-training-dataset Bitext - Media Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [media] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-media-llm-chatbot-training-dataset.textquestion-answering10K<n<100K0 likes117 downloads2y agoHugging Face27bitext /Bitext-restaurants-llm-chatbot-training-dataset Bitext - Restaurants Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [restaurants] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-restaurants-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes110 downloads2y agoHugging Face28noob123 /imdb_traintext1K<n<10K0 likes107 downloads4y agoHugging Face29sschet /ROCOv2-traintext10K<n<100K0 likes107 downloads2y agoHugging Face30CreitinGameplays /DeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1tabular1K<n<10K0 likes101 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.