Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01manycore-research /SpatialLM-Testset SpatialLM Testset Project page | Paper | Code We provide a test set of 107 preprocessed point clouds and their corresponding GT layouts, point clouds are reconstructed from RGB videos using MASt3R-SLAM. SpatialLM-Testset is quite challenging compared to prior clean RGBD scan datasets due to the noises and occlusions in the point clouds reconstructed from monocular RGB videos. Folder Structure Outlines of the dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Testset.3dn<1K60 likes1.2k downloads1y agoHugging Face02sunday-hao /vindr-cxr-testsetimage1K<n<10K0 likes838 downloads3mo agoHugging Face03allegrolab /testset_piqatext1K<n<10K0 likes303 downloads1y agoHugging Face04allegrolab /testset_popqatext1K<n<10K0 likes242 downloads1y agoHugging Face05allegrolab /testset_mmlutext1K<n<10K0 likes237 downloads1y agoHugging Face06allegrolab /testset_winogrande-infilltext1K<n<10K0 likes221 downloads1y agoHugging Face07allegrolab /testset_hellaswagtext1K<n<10K0 likes207 downloads1y agoHugging Face08KinGeorge /Dr.Sparse-OTF-test-set Dr.Sparse OTF Test Set 100 sparse matrices from the SuiteSparse Matrix Collection, converted to the flat binary format the Dr.Sparse benchmark harness reads. This is the held-out evaluation set for LLM-generated CUDA sparse kernels (SpMV / SpMM / SpGEMM), kept separate from the matrices the models were developed against. Layout Matrices are grouped into size tiers by row count, the convention Dr.Sparse task discovery scans for: tier rows matrices size… See the full description on the dataset page: https://huggingface.co/datasets/KinGeorge/Dr.Sparse-OTF-test-set.tabularothern<1K0 likes93 downloads1mo agoHugging Face09allegrolab /testset_winogrande-mcqtext1K<n<10K0 likes59 downloads1y agoHugging Face10OpenVoiceOS /ovos-intents-ilenia-testset-nl Retired. This dataset is superseded by OpenVoiceOS/ovos-intents. It stays available for reproducibility and receives no updates. texttext-classificationn<1K0 likes39 downloads6d agoHugging Face11omarsou /common_voice_16_1_spanish_test_set Dataset Card for Common Voice Corpus 16 Spanish Dataset Acknowledgement The dataset belongs to COMMON VOICE MOZILLA FOUNDATION. I just uploaded the spanish test set (from HERE : https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1/tree/main) Dataset Summary The Common Voice dataset consists of a unique MP3 and corresponding text file. Languages Spanish How to use The datasets library allows you to load and pre-process… See the full description on the dataset page: https://huggingface.co/datasets/omarsou/common_voice_16_1_spanish_test_set.tabular10K<n<100K1 likes34 downloads3y agoHugging Face12allegrolab /testset_munchtextn<1K0 likes34 downloads1y agoHugging Face13OpenVoiceOS /ovos-intents-ilenia-testset-es Retired. This dataset is superseded by OpenVoiceOS/ovos-intents. It stays available for reproducibility and receives no updates. texttext-classificationn<1K0 likes33 downloads6d agoHugging Face14allegrolab /testset_ellietextn<1K0 likes31 downloads1y agoHugging Face15OpenVoiceOS /ovos-intents-ilenia-testset-ca Retired. This dataset is superseded by OpenVoiceOS/ovos-intents. It stays available for reproducibility and receives no updates. texttext-classificationn<1K0 likes31 downloads6d agoHugging Face16Tarive /within_family_test_settext100K<n<1M0 likes31 downloads1y agoHugging Face17parnoux /hate_speech_open_data_original_class_test_settabulartext-classification1K<n<10K1 likes30 downloads4y agoHugging Face18zhangyingbo1984 /Pharmacology-LLM-test-setPharmacology-LLM-test-set: A test set for a large language model focused on pharmacology tasks 1 Inroduction Large language models (LLM), including ChatGPT, have fundamentally transformed the knowledge query schemes and methods in pharmacology for pharmacologists, drug researchers, clinical drug researchers, and artificial intelligence researchers in pharmacology. They can conduct multi-round consultations and query pharmacological issues in a question-and-answer format. However… See the full description on the dataset page: https://huggingface.co/datasets/zhangyingbo1984/Pharmacology-LLM-test-set.textn<1K2 likes29 downloads2y agoHugging Face19AlexWu /Vox2_testsettabularn<1K0 likes29 downloads3mo agoHugging Face20yiran223 /toxic-detection-testset-perturbations Dataset Card for toxic-detection-testset-perturnations Dataset Summary This dataset a test set for toxic detection that contains both clean data and it's perturbed version with human-written perturbations online. In addition, our dataset can be used to benchmark misspelling correctors as well. Supported Tasks and Leaderboards [More Information Needed] Languages English Dataset Structure Data Instances { "clean_version": "this… See the full description on the dataset page: https://huggingface.co/datasets/yiran223/toxic-detection-testset-perturbations.tabular1K<n<10K0 likes22 downloads4y agoHugging Face21seregadgl /test_settext100K<n<1M0 likes16 downloads5y agoHugging Face22prasannad28 /augmented_posts_test_settextsentence-similarity1K<n<10K0 likes14 downloads2y agoHugging Face23ianyang02 /aita_test_set_balanced_4_7_26text1K<n<10K0 likes11 downloads6mo agoHugging Face24climate-adaptation /adaptation-test-setThis is the training data for the models of the project. tabular1K<n<10K0 likes11 downloads5mo agoHugging Face25Elkelouizajo /hans_reduced_testsettabular1K<n<10K0 likes8 downloads2y agoHugging Face26prasannad28 /translated_facts_test_settext1K<n<10K0 likes8 downloads2y agoHugging Face27Chaksseu /mmg_clotho_test_settext1K<n<10K0 likes8 downloads2y agoHugging Face28MoGP /g_test_set_newtext10K<n<100K0 likes7 downloads2y agoHugging Face29sahask8 /cb_test_settextn<1K0 likes7 downloads2y agoHugging Face30johntzwei /testset_winogrande-mcqtext1K<n<10K0 likes7 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.