datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imagefolder_with_metadatafixtures_ade20kfixtures-cocotest-videostokenizers-benchexample-imageszen-imageimagefolder_with_metadata_no_splitstransformers-synthetic-assets
transformers-synthetic-assets
Synthetic media fixtures for Transformers tests. These assets are generated from prompts or deterministic code and are not derived from third-party source files.
dummy_image_text_data
Dataset Card for "dummy_image_text_data"
More Information needed
zen-multi-imagefixtures_docvqaThis dataset includes 2 document images of the DocVQA dataset.
They are used for testing the LayoutLMv2FeatureExtractor + LayoutLMv2Processor inside the HuggingFace Transformers library.
More specifically, they are used in tests/test_feature_extraction_layoutlmv2.py and tests/test_processor_layoutlmv2.py.
fill10cats_vs_dogs_samplesam2-fixturesimages_testAdaptFM_TestingBBBC024
3D HL60 Cell line (synthetic data). This dataset contains a synthetic images of 3D cell nuclei. Images are stratified by clustering probability and SNR (low vs high). A complete description of the dataset can be found here
BBBC027
3D Colon tissue synthetic images. This dataset has synthentic images of 3D colon organoids. Images are stratified by low vs high SNR. A full description of the dataset can be found here
RESECT
Intraoperative brain tumor ultrasound images before, during, and… See the full description on the dataset page: https://huggingface.co/datasets/hbakhtiar/AdaptFM_Testing.dummy_image_class_data
Dataset Card for "dummy_image_class_data"
More Information needed
testing-imagespeft-blog-assetsAssets for PEFT blog posts
mask-for-image-segmentation-testsfixtures_ocrThis dataset includes 2 images: one of the IAM Handwriting Database and one of the SRIOE dataset.
They are used for testing OCR models that are part of the HuggingFace Transformers library. See here for details.
More specifically, they are used inside test_modeling_vision_encoder_decoder_model.py, for testing the TrOCR models.
controlnet-testinginstructpix2pix-10-samples
Dataset Card for "test"
More Information needed
so101_tb4_RJ_testing_RB0fixtures_got_ocrdocument-visual-retrieval-test
Model Card: Document Visual Retrieval Test (internal)
Dataset Overview
This dataset is designed to evaluate the performance of visual retrievers by testing their ability to match a query to a relevant image. Each of the three examples in this dataset contains a text query and an associated image, which is a scanned page from the foundational "Attention is All You Need" paper. The purpose of this dataset is to facilitate the evaluation of visual retrievers, where the… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/document-visual-retrieval-test.image-matting-fixturescolmap-testing-dataset
COLMAP Testing Dataset
A small, ready-to-run collection of multi-view image scenes for testing and
benchmarking the COLMAP Structure-from-Motion (SfM)
and Multi-View Stereo (MVS) pipeline — used here to validate a macOS / Apple
Silicon (Metal) build of COLMAP, but useful for any SfM/MVS work.
It bundles two well-known families of scenes, each already laid out in the
directory structure COLMAP expects, so you can point COLMAP at a folder and run
the pipeline end-to-end with no… See the full description on the dataset page: https://huggingface.co/datasets/alexmkwizu/colmap-testing-dataset.testing-resources
