Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01stablellama /Krea-2-Raw_samplesThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw. NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model. Raw is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on Krea 2 Raw Quality testing Data source The images were created in ComfyUI with the bf16 version of… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples.tabulartext-to-image1K<n<10K0 likes2k downloads1mo agoHugging Face02SBMM75 /Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw. NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model. Raw is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on Krea 2 Raw Quality testing Data source This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/SBMM75/Krea-2-Raw_samples_Best_of.tabulartext-to-image1K<n<10K0 likes961 downloads28d agoHugging Face03stablellama /Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw. NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model. Raw is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on Krea 2 Raw Quality testing Data source This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples_Best_of.tabulartext-to-image1K<n<10K0 likes471 downloads1mo agoHugging Face04MUG-V /MUG-V-Training-Samples MUG-V Training Samples Sample training dataset for the MUG-V 10B video generation model training framework. Dataset Description This dataset contains pre-processed training samples for quick-start validation and testing of the MUG-V Megatron-LM training pipeline. It includes: VideoVAE-encoded latents (8×8×8 compressed video representations) T5-XXL text features (4096-dim embeddings) Training metadata CSV (sample mapping and configuration) ⚠️ Note: This is a sample… See the full description on the dataset page: https://huggingface.co/datasets/MUG-V/MUG-V-Training-Samples.texttext-to-video1K<n<10K0 likes322 downloads1y agoHugging Face05Lab-Rasool /honeybee-samples HoneyBee Sample Files Sample data and resource files for the HoneyBee framework — a scalable, modular toolkit for multimodal AI in oncology. These files are used by the HoneyBee example notebooks (clinical, pathology, radiology) and by HoneyBee's molecular processing code at runtime (Hugo_symbols.tsv is fetched on first use of DNA mutation preprocessing). Paper: HoneyBee: A Scalable Modular Framework for Creating Multimodal Oncology Datasets with Foundational Embedding Models… See the full description on the dataset page: https://huggingface.co/datasets/Lab-Rasool/honeybee-samples.documentimage-classification10K<n<100K1 likes104 downloads5mo agoHugging Face06locationlists /us-business-locations-samples LocationLists — US business location samples Ten real rows from each of 30 of the 726 datasets published at locationlists.com, a catalog of 43,101,581 verified US business locations: manufacturer dealer networks, retail and restaurant chains, licensed trade contractors, healthcare providers, nonprofits and membership directories. Every dataset is compiled from the brand's own store locator or the official government registry, re-checked weekly, and sold as a flat CSV with no… See the full description on the dataset page: https://huggingface.co/datasets/locationlists/us-business-locations-samples.tabularn<1K0 likes82 downloads27d agoHugging Face07maslennikovds /alion-job-market-samples Alion job-market data samples Free samples of the datasets behind Alion: job postings read directly from employers' own applicant tracking systems and career pages, the companies behind them, and pay benchmarks computed from those postings. The same files are downloadable from https://alion.io/data. Config Rows What a row is job_postings 2,500 One opening read from an employer's ATS board or career page: role and role family, seniority, employment type, work mode… See the full description on the dataset page: https://huggingface.co/datasets/maslennikovds/alion-job-market-samples.tabular1K<n<10K1 likes73 downloads17d agoHugging Face08Parthsoni10 /supplychain-agent-input-samples Supply Chain Agent Input Samples Synthetic input samples for a five-agent supply-chain platform (aizenio/supplychain-ops). Each row is one internally consistent snapshot that satisfies the input contract of every agent skill — 67 columns covering demand forecasting, inventory monitoring, logistics/shipment evaluation, anomaly detection and strategic analysis. A single row can be fed to any agent without post-processing. Why it exists The agents needed realistic… See the full description on the dataset page: https://huggingface.co/datasets/Parthsoni10/supplychain-agent-input-samples.tabulartabular-classification100K<n<1M0 likes68 downloads29d agoHugging Face09gmreincglm /usta-feeds-samples Dated samples of United States public-record change files 15 samples, one folder per family. Each folder holds sample.csv, sample.json and a README naming the source, the columns, the sealing date and the row count. Every file is a change file, not a snapshot. We seal dated copies of a public source, compare two copies, and keep what appeared, what stopped being listed, and what quietly changed in between. Most of these sources publish only the list as it stands today and… See the full description on the dataset page: https://huggingface.co/datasets/gmreincglm/usta-feeds-samples.tabularn<1K0 likes61 downloads1mo agoHugging Face10bibekyess /layout-detector-flagged-samples Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/bibekyess/layout-detector-flagged-samples.imagen<1K0 likes43 downloads3y agoHugging Face11Luel-ai /luel-multilingual-tts-samplesgated Multilingual TTS Samples (Luel) License: All Rights Reserved. Proprietary. Access only for authorized parties; no redistribution or use without permission. See LICENSE. A multilingual text-to-speech / read-speech dataset of short scripted utterances across 7 languages. Each sample is a single-speaker recording of a written prompt, paired with rich speaker and recording metadata. Useful for TTS training and evaluation, ASR adaptation, dialect/accent studies, and read-speech… See the full description on the dataset page: https://huggingface.co/datasets/Luel-ai/luel-multilingual-tts-samples.audiotext-to-speechn<1K0 likes42 downloads5mo agoHugging Face12baber /pd_books_samplestextn<1K0 likes36 downloads2y agoHugging Face13helvia /banking77-representative-samples About This is a curated subset of 3 representative samples per class (77 classes in total) for the Banking77 dataset, as collected by a domain expert. It was used in the paper "Making LLMs Worth Every Penny: Resource-Limited Text Classification in Banking", published in ACM ICAIF 2023 (https://arxiv.org/abs/2311.06102). Our findings show that Few-Shot Text Classification on representative samples are better than randomly selected samples. Citation… See the full description on the dataset page: https://huggingface.co/datasets/helvia/banking77-representative-samples.textn<1K2 likes34 downloads3y agoHugging Face14andrewy1n /pseudocode-decompiled-samples-smalltext1K<n<10K0 likes30 downloads3y agoHugging Face15genbio-ai /sample-structure-dataset Sample dataset for PETAL model This dataset is a sample dataset to test the functionalities of the PETAL model (encoder and decoder). It is based on CASP15 dataset, see https://predictioncenter.org/casp15/ https://github.com/Bhattacharya-Lab/CASP15 The registries folder contains the registry of CASP15 dataset (a csv file with filename, pdb_id, etc.) tabularn<1K0 likes27 downloads2y agoHugging Face16bryanchrist /EDUMATH_model_samples Math Word Problems from Comparison Models in EDUMATH: Generating Standards-aligned Educational Math Word Problems This dataset contains 8,360 math word problems annotated by Gemma 3 27B IT and the EDUMATH Classifier from the models compared in EDUMATH: Generating Standards-aligned Educational Math Word Problems. Each row contains a question and answer along with the grade level and math standard(s) it was generated for and the model it was generated from. The Gemma 3 27B IT label… See the full description on the dataset page: https://huggingface.co/datasets/bryanchrist/EDUMATH_model_samples.tabular1K<n<10K0 likes27 downloads6mo agoHugging Face17h1alexbel /github-samplestext1K<n<10K1 likes24 downloads2y agoHugging Face18hoorangyee /biglawbench_reversed_score_samplestextn<1K0 likes19 downloads2y agoHugging Face19jozhang97 /improved-genie2-samplestabularn<1K0 likes17 downloads1y agoHugging Face20Svetlana0303 /all_samples_Regressiontabularn<1K1 likes15 downloads4y agoHugging Face21bhargavi909 /mt_Samplestext1K<n<10K1 likes13 downloads3y agoHugging Face22Anonymous-EmpathyAI /300-samples-cof-reftextn<1K0 likes13 downloads1y agoHugging Face23jozhang97 /ambient-long-samplestabularn<1K0 likes11 downloads1y agoHugging Face24jozhang97 /ambient-short-samplestabular1K<n<10K0 likes11 downloads1y agoHugging Face25motionlabs /cosmopedia-v2-ranked-samplestext100K<n<1M0 likes11 downloads1y agoHugging Face26vmodak /rag-eval-100-samplestextn<1K0 likes11 downloads1y agoHugging Face27Soumyajitxedu /sample_students Dataset Card for Holyfaith Academy Student Records This dataset contains academic records for Class VIII and VII students. Personal Identifiable Information (PII) like phone numbers and birth years has been masked to ensure privacy. Data Fields Field Name Description CARD NO Unique identification number for the student record. Name Full name of the student in uppercase. OLD Class Academic level (e.g., VIII or VII). Old.UNIT The school unit or section… See the full description on the dataset page: https://huggingface.co/datasets/Soumyajitxedu/sample_students.tabulartoken-classificationn<1K0 likes11 downloads10mo agoHugging Face28asr-malayalam /Norm_Malayalam_Evaluation_samples Malayalam ASR Reference Prediction dataset This repository contains evaluation results from the Malayalam ASR model "vrclc/Whisper_small_malayalam" using the "google/fleurs" dataset. ASR Model Name: vrclc/Whisper_small_malayalam Dataset: google/fleurs Curated by: VRCLC vrclc/Whisper_small_malayalam was trained with 50 hours of Malayalam speech data. The test set of google/fleurs dataset which consists of Malayalam speech data was used to evaluate the model The evaluation of 500… See the full description on the dataset page: https://huggingface.co/datasets/asr-malayalam/Norm_Malayalam_Evaluation_samples.textsentence-similarityn<1K0 likes8 downloads2y agoHugging Face29laion /tts-samples-with-artifactstabular1K<n<10K0 likes8 downloads4mo agoHugging Face30chjlhxww /final_samples_for_labeling_LongContexttext100K<n<1M0 likes7 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.