datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
transformers_circleci_workflow_runsdataset_with_scriptThis is a test dataset.audiofolder_two_configs_in_metadatalibrispeech_asr_dummyaudiofolder_single_config_in_metadatavideos-testzenimagefolder_with_metadatadataset_with_data_filesmulti_dir_datasetdiffusers-imagesaudiofolder_no_configs_in_metadatatokenizers-test-data
tokenizers-test-data
Test and benchmark fixtures for huggingface/tokenizers,
pulled on demand by the repo Makefiles (make test / make bench / make fixtures
via hf download).
Layout
fixtures/ — multilingual + modality corpora for cross-language encode
benchmarks. Organized, documented, and reproducible: see
fixtures/FIXTURES.md for provenance and
fixtures/fixtures_manifest.json for
exact sources, pinned revisions, and sizes. Rebuild any file with… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/tokenizers-test-data.DatasetWithCapitalLettersfixtures_image_utils\\nraw_jsonlfixtures_ade20kfixtures-cocotesting_alpaca_small
Dataset Card for "testing_alpaca_small"
More Information needed
audiofolder_two_configs_in_metadata_with_defaulttest-videosdummy-audio-samplestesting_self_instruct_small
Dataset Card for "testing_self_instruct_small"
More Information needed
tokenizers-benchzen-imagehfh_ci_scan_dataset_bexample-imagesimagefolder_with_metadata_no_splitscompressed_filestransformers_flash_attn_ci
