Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HCOOH /vlm_wsi_patch VLM WSI Patch Closed-answer WSI patch VQA dataset with 48,766 metadata-eligible rows and 48,609 valid 1344×1344 RGB top-4 mosaics, picked from of PathGen. Rows without an image retain an explicit processing rejection status. Each patch uses a WSI level-0 top-left coordinate and 672×672 pixels; backfill and patch substitution are excluded. The image field is a repository-relative PNG path. metadata/manifest.jsonl and metadata/manifest.parquet contain QA labels, top-4 coordinates… See the full description on the dataset page: https://huggingface.co/datasets/HCOOH/vlm_wsi_patch.imagevisual-question-answering0 likes2.4k downloads3mo agoHugging Face02federicosabbadini /patch-aliasing-bayesian Patch-Aliasing Bayesian Analysis Data Raw and processed data from the Bayesian analysis of structural patch-aliasing in Chronos-Bolt Tiny. Companion to the model weights at federicosabbadini/chronos-bolt-patch-aliasing-models. Structure clean_15model/ # 15 preregistered (P,S) configurations full_22model/ # all 22 configurations (adds 7 robustness checks) h1_fixed_offset/ # supplementary H1 analysis with fixed offset Each run folder… See the full description on the dataset page: https://huggingface.co/datasets/federicosabbadini/patch-aliasing-bayesian.imagen<1K1 likes751 downloads2mo agoHugging Face03anaumghori /patchlet-embed-preprocessedimage100K<n<1M0 likes540 downloads8mo agoHugging Face04PatchMap /PatchMap_v1image1K<n<10K0 likes501 downloads1y agoHugging Face05closji /cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13image10M<n<100M0 likes496 downloads4y agoHugging Face06CosVersin /e621-tagger-patchimagen<1K3 likes478 downloads3y agoHugging Face07dpdl-benchmark /patch_camelyonimage100K<n<1M0 likes292 downloads2y agoHugging Face08DefendIntelligence /vessel-detection-labeled-patches Vessel Detection Labeled Patches Validated/confirmed satellite image patches exported from the military-boat-detection review workflow. Contents images/: patch images. metadata.csv: one row per patch, compatible with Hugging Face image-folder metadata. metadata.jsonl: rich patch metadata with nested objects. annotations.csv: one row per vessel annotation. annotations.jsonl: JSONL version of the object annotations. labels/: YOLO-format labels. Hard negatives have empty… See the full description on the dataset page: https://huggingface.co/datasets/DefendIntelligence/vessel-detection-labeled-patches.imageobject-detection1K<n<10K5 likes292 downloads5mo agoHugging Face09zacharielegault /PatchCamelyon PatchCamelyon (PCam) This is a reupload of the PatchCamelyon (PCam) dataset to make it more readily usable instead of manipulating H5 files. The original can be found in the author's Github repo. If you use this dataset, please cite the original publications: @inproceedings{veeling2018rotation, title={Rotation Equivariant CNNs for Digital Pathology}, author={Veeling, Bastiaan S and Linmans, Jasper and Winkens, Jim and Cohen, Taco and Welling, Max}, booktitle={Medical Image… See the full description on the dataset page: https://huggingface.co/datasets/zacharielegault/PatchCamelyon.imageimage-classification100K<n<1M2 likes278 downloads2y agoHugging Face10closji /cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15image10M<n<100M0 likes196 downloads4y agoHugging Face11YoussefMoNader /ink-8um-v8-patchpack ink-8um v8 patch pack This is the complete training and validation corpus of the v8-in ink-detection model, frozen to disk as 64 × 64 × 24 uint8 patches. It holds 508,006 training and 43,249 validation patches from 19 surface segments at ~8 µm: Scroll 1, Scroll 5, PHerc. 1667, 0139, 0814, 0500P2 and 0009B, plus fragments 1, 3 and 4. The patches were extracted with the training code's own loaders and admission rules, so training from the pack is equivalent to training from the… See the full description on the dataset page: https://huggingface.co/datasets/YoussefMoNader/ink-8um-v8-patchpack.imagen<1K0 likes151 downloads9d agoHugging Face12EastDong0704 /patch-of-behavior-1k BEHAVIOR-1K Asset Database (Patch) Complete metadata and visualizations for all 8662 BEHAVIOR-1K object assets. Contents Path Description assets.db SQLite database — descriptions + DoF for all 8662 assets objects/{category}/{model}/description.json VLM-generated visual description (5676 assets) objects/{category}/{model}/dof.json Articulation / DoF metadata objects/{category}/{model}/front.png Front-view visualization objects/{category}/{model}/back.png… See the full description on the dataset page: https://huggingface.co/datasets/EastDong0704/patch-of-behavior-1k.imageimage-to-3d10K<n<100K0 likes134 downloads7mo agoHugging Face13letitiaaa /patch_region_128image100K<n<1M0 likes133 downloads9mo agoHugging Face14TrevorJS /mtg-scryfall-cropped-art-embeddings-siglip-so400m-patch14-384image10K<n<100K0 likes127 downloads2y agoHugging Face15closji /cc12m_openai-clip-vit-patch32image1M<n<10M1 likes123 downloads4y agoHugging Face16bramtoula /vtab_patch_camelyon VTAB PatchCamelyon This dataset has been used for the paper Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models (NeurIPS 2025). It reproduces the settings (splits, labels) used for the Visual Task Adaptation Benchmark (VTAB). VTAB Paper: A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark VTAB Repository: google-research/task_adaptation Details of the original dataset: Original… See the full description on the dataset page: https://huggingface.co/datasets/bramtoula/vtab_patch_camelyon.image100K<n<1M0 likes108 downloads7mo agoHugging Face17letitiaaa /patch_region_512image100K<n<1M0 likes102 downloads9mo agoHugging Face18pavan316 /PatchCamelyon PatchCamelyon (PCam) Description The PatchCamelyon benchmark is a new and challenging image classification dataset. It consists of 327.680 color images (96 x 96px) extracted from histopathologic scans of lymph node sections. Each image is annoted with a binary label indicating presence of metastatic tissue. PCam provides a new benchmark for machine learning models: bigger than CIFAR10, smaller than imagenet, trainable on a single GPU Why PCam Fundamental… See the full description on the dataset page: https://huggingface.co/datasets/pavan316/PatchCamelyon.imageimage-classification100K<n<1M0 likes87 downloads9mo agoHugging Face19ssuresh /idc-patchesimage100K<n<1M0 likes85 downloads11mo agoHugging Face20KE9037 /PatchCamelyon PatchCamelyon (PCam) Description The PatchCamelyon benchmark is a new and challenging image classification dataset. It consists of 327.680 color images (96 x 96px) extracted from histopathologic scans of lymph node sections. Each image is annoted with a binary label indicating presence of metastatic tissue. PCam provides a new benchmark for machine learning models: bigger than CIFAR10, smaller than imagenet, trainable on a single GPU Why PCam Fundamental… See the full description on the dataset page: https://huggingface.co/datasets/KE9037/PatchCamelyon.imageimage-classification100K<n<1M0 likes79 downloads10mo agoHugging Face21closji /cc12m_openai-clip-vit-patch32_image_retrieval_top15_start1000000_end3500000image1M<n<10M0 likes75 downloads4y agoHugging Face22ASTRALK /sentinel-lfm-mining-patches sentinel-lfm — illegal-mining single-frame patches 128px RGB patches cropped from the Roboflow illegal-mining dataset, labelled mine (1) / no-mine (0). Split by source image (no leakage) into train/val/test. Provided as PNGs + vlm_sft-format JSONL (one image + prompt -> JSON answer) so it drops straight into VLM fine-tuning. split pos neg total train 1410 555 1965 val 303 66 369 test 303 116 419 RGB only (no multispectral). Each JSONL row is a single-turn VLM… See the full description on the dataset page: https://huggingface.co/datasets/ASTRALK/sentinel-lfm-mining-patches.imageimage-classification1K<n<10K0 likes73 downloads4mo agoHugging Face23ykeselman /lidc-idri-patchesA dataset of patches (most 64x64 pixels from the LIDC-IDRI CT scans). These are 16-bit images with a 1024 shift from the original HU values. The dataset is incomplete as of 2024-12-02. If you find it useful, I will add more patches. --- license: apache-2.0 task_categories: - image-classification language: - en pretty_name: LIDC IDRI patches for classification configs: - config_name: default data_files: - split: train path: data/train-* dataset_info: features: - name: annotation_id… See the full description on the dataset page: https://huggingface.co/datasets/ykeselman/lidc-idri-patches.image10K<n<100K1 likes71 downloads2y agoHugging Face24AngeloUNIMI /ALL-IDB-Patches ALL-IDB Patches MATLAB source code for creating image patches and labels used in the paper “ALL-IDB Patches: Whole slide imaging for Acute Lymphoblastic Leukemia detection using Deep Learning”, presented at ICASSP Workshops 2023. The repository converts annotated ALL-IDB1 whole-slide microscope images into fixed-size overlapping patches, preserving the position of white blood cell centroids and generating patch-level labels for probable lymphoblasts… See the full description on the dataset page: https://huggingface.co/datasets/AngeloUNIMI/ALL-IDB-Patches.imageimage-classificationn<1K0 likes64 downloads4mo agoHugging Face25jonathan-roberts1 /Airbus-Wind-Turbines-Patches Dataset Card for "Airbus-Wind-Turbines-Patches" Split Information This HuggingFace dataset repository contains just the Validation split. Licensing Information CC BY-NC-SA 4.0 Citation Information Airbus Wind Turbine Patches @misc{kaggle_awtp, author = {Airbus DS GEO S.A.}, title = {Airbus Wind Turbine Patches}, howpublished = {\url{https://www.kaggle.com/datasets/airbusgeo/airbus-wind-turbines-patches}}, year = {2021}, version = {1.0} } image10K<n<100K1 likes47 downloads4y agoHugging Face26yurkes /patch_tasks_vllm Dataset Card for Patch-Based Visual Question Answering Dataset Dataset Details Dataset Description This dataset contains approximately 305,000 triplets of question, answer, and image designed for patch-based visual reasoning tasks. A standard question in this dataset is formatted as follows: Image Grid: The image is divided into a 4x4 grid of 16 equal-sized patches. Patches are numbered sequentially from the top-left corner and moving right, then down to the… See the full description on the dataset page: https://huggingface.co/datasets/yurkes/patch_tasks_vllm.imageimage-text-to-text100K<n<1M3 likes39 downloads1y agoHugging Face27ducido /merged_libero_scale_40_mask_depth_noops_lerobot_patch_size7_chunk10_mask_interestimage100K<n<1M0 likes37 downloads9mo agoHugging Face28closji /cc12m_openai-clip-vit-patch32_image_retrieval_top15_start1000000_end3500000_SHORT500Kimage100K<n<1M0 likes35 downloads4y agoHugging Face29jlee124 /A85_patch_52imagen<1K1 likes34 downloads1y agoHugging Face30pykale /bln600-img-patch BLN600 Image Patches This dataset provides BLN600's image patches for fine-tuning vision-language models on post-OCR correction, introduced in "Image-Informed Post-OCR Correction with Vision-Language Models" (EMNLP 2026 Findings). Each patch corresponds to a text sequence in BLN600 and is cropped from the Gale British Library Newspapers collection, using word-level bounding boxes from the collection's ALTO XML OCR layout data. It is intended to be used alongside the code and… See the full description on the dataset page: https://huggingface.co/datasets/pykale/bln600-img-patch.imageimage-to-text10K<n<100K0 likes33 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.