Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-source-metrics /tokenizers-dependents tokenizers metrics This dataset contains metrics about the huggingface/tokenizers package. Number of repositories in the dataset: 11460 Number of packages in the dataset: 124 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 14 packages that have more than 1000 stars. There are 41… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/tokenizers-dependents.tabularn<1K0 likes1.2k downloads2y agoHugging Face02open-source-metrics /transformers-dependents transformers metrics This dataset contains metrics about the huggingface/transformers package. Number of repositories in the dataset: 27067 Number of packages in the dataset: 823 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 65 packages that have more than 1000 stars. There are 140… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/transformers-dependents.tabular10K<n<100K2 likes976 downloads2y agoHugging Face03HKUSTAudio /Llasa_opensource_speech_data_160k_hours_tokenized Update (2025-02-07): Our paper has been released! This script is for merging tokenized speech datasets stored in memmap format. The input datasets can be combined to form larger training datasets. import numpy as np import os def merge_memmap_datasets(dataset_dirs, output_dir): # Ensure the output directory exists os.makedirs(output_dir, exist_ok=True) # Dataset splits to be merged splits = ['train', 'val'] for split in splits: shapes = [] seq_len =… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/Llasa_opensource_speech_data_160k_hours_tokenized.31 likes808 downloads2y agoHugging Face04open-source-metrics /evaluate-dependents evaluate metrics This dataset contains metrics about the huggingface/evaluate package. Number of repositories in the dataset: 106 Number of packages in the dataset: 3 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 1 packages that have more than 1000 stars. There are 2 repositories… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/evaluate-dependents.tabular1K<n<10K0 likes784 downloads2y agoHugging Face05open-source-metrics /gradio-dependents Dataset Card for "gradio-dependents" More Information needed tabular1K<n<10K0 likes780 downloads2y agoHugging Face06open-source-metrics /diffusers-dependents diffusers metrics This dataset contains metrics about the huggingface/diffusers package. Number of repositories in the dataset: 160 Number of packages in the dataset: 2 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 0 packages that have more than 1000 stars. There are 3 repositories… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/diffusers-dependents.tabular1K<n<10K1 likes695 downloads2y agoHugging Face07open-source-metrics /optimum-dependents optimum metrics This dataset contains metrics about the huggingface/optimum package. Number of repositories in the dataset: 19 Number of packages in the dataset: 6 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 0 packages that have more than 1000 stars. There are 0 repositories that… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/optimum-dependents.tabularn<1K1 likes551 downloads2y agoHugging Face08open-source-metrics /accelerate-dependents accelerate metrics This dataset contains metrics about the huggingface/accelerate package. Number of repositories in the dataset: 727 Number of packages in the dataset: 37 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 10 packages that have more than 1000 stars. There are 16… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/accelerate-dependents.tabular1K<n<10K1 likes543 downloads2y agoHugging Face09open-source-metrics /datasets-dependents datasets metrics This dataset contains metrics about the huggingface/datasets package. Number of repositories in the dataset: 4997 Number of packages in the dataset: 215 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 22 packages that have more than 1000 stars. There are 43… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/datasets-dependents.tabular10K<n<100K0 likes511 downloads2y agoHugging Face10open-source-metrics /pip Dataset Card for "pip" More Information needed text10K<n<100K0 likes345 downloads2y agoHugging Face11simbahuang /wan22-animate-3k-opensource-data Wan2.2 Animate Open Dataset Pack This dataset repo stores the complete datasets/ directory used for the Wan2.2 TI2V 5B + One-to-All animate experiment. The original tree contains more than 10,000 files in one directory, which Hugging Face git repositories reject as raw files. Therefore the dataset is stored as split tar shards. Restore: cat datasets.tar.part-* | tar -xf - sha256sum -c SHA256SUMS After extraction, the restored tree contains:… See the full description on the dataset page: https://huggingface.co/datasets/simbahuang/wan22-animate-3k-opensource-data.image10K<n<100K0 likes284 downloads3mo agoHugging Face12open-source-metrics /pytorch-image-models-dependents pytorch-image-models metrics This dataset contains metrics about the huggingface/pytorch-image-models package. Number of repositories in the dataset: 3615 Number of packages in the dataset: 89 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 18 packages that have more than 1000… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/pytorch-image-models-dependents.1 likes248 downloads2y agoHugging Face13TongueProjectTCM /color-classification-opensourceimagen<1K1 likes220 downloads2y agoHugging Face14open-source-metrics /starstext100K<n<1M0 likes217 downloads2y agoHugging Face15open-source-metrics /issuesimage100K<n<1M0 likes169 downloads2y agoHugging Face16open-source-metrics /reinforcement-learning-checkpoint-downloadstextn<1K5 likes160 downloads4y agoHugging Face17gemmozero /ai-opensource-2026gated Ai Opensource 2026 Part of the LEGION Intelligence dataset collection. Provider: LEGION Systems Access: Requires approval — submit request below Usage from datasets import load_dataset dataset = load_dataset("gemmozero/ai-opensource-2026") API Access Real-time access via LEGION API: curl https://api.legion-api.com/incidents API Docs · Pro Access €29/mo License CC BY-NC 4.0 — Research and non-commercial use only. Commercial use… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-opensource-2026.texttext-classificationn<1K0 likes158 downloads5d agoHugging Face18HansBug /opensource_mirror0 likes155 downloads3y agoHugging Face19andyP /fake_news_en_opensources Dataset Card for "Fake News Opensources" Dataset Description Homepage: https://github.com/AndyTheFactory/FakeNewsDataset Repository: https://github.com/AndyTheFactory/FakeNewsDataset Point of Contact: Andrei Paraschiv Dataset Summary a consolidated and cleaned up version of the opensources Fake News dataset Fake News Corpus comprises 8,529,090 individual articles, classified into 12 classes: reliable, unreliable, political, bias, fake, conspiracy… See the full description on the dataset page: https://huggingface.co/datasets/andyP/fake_news_en_opensources.texttext-classification1M<n<10M2 likes145 downloads3y agoHugging Face20open-source-benchmarking /os-world-modified0 likes138 downloads1y agoHugging Face21Agatha7k /opensource_100_TVG_casevideon<1K0 likes133 downloads1y agoHugging Face22open-source-metrics /visual-question-answering-checkpoint-downloadstabularn<1K7 likes126 downloads4y agoHugging Face23open-source-metrics /hub-docs-dependents Dataset Card for "hub-docs-dependents" More Information needed 0 likes110 downloads2y agoHugging Face24open-source-metrics /safetensors-dependents Dataset Card for "safetensors-dependents" More Information needed tabular1K<n<10K0 likes108 downloads2y agoHugging Face25Cyber-security-final-project /Evaluation_of_OpenSource_Models_for_PDF_Injection_Recognition Injected PDFs - Model Evaluation This repository holds the model evaluation stage of a project on detecting harmless-but-real attack payloads injected into PDF files, together with the artefacts it produced for the application. Nothing is trained here. Seven off-the-shelf models are measured against the same 1,100 PDFs, and the two winners are exported for the app to load. Question Candidates Winner Part A Which files look like this one? 3 embedding models x 2 inputs… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/Evaluation_of_OpenSource_Models_for_PDF_Injection_Recognition.tabulartext-classification1K<n<10K0 likes107 downloads2mo agoHugging Face26softcatala /open-source-english-catalan-corpus Dataset Card for open-source-english-catalan-corpus Dataset Summary Translation memory built from more than 180 open source projects. These include LibreOffice, Mozilla, KDE, GNOME, GIMP, Inkscape and many others. It can be used as translation memory or as training corpus for neural translators. Supported Tasks and Leaderboards [More Information Needed] Languages Catalan (ca) English (en) Dataset Structure Data Instances [More… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/open-source-english-catalan-corpus.texttext-generationn<1K1 likes101 downloads4y agoHugging Face27open-source-metrics /unconditional-image-generation-checkpoint-downloadstextn<1K4 likes99 downloads4y agoHugging Face28open-source-metrics /issues-externaltext1M<n<10M0 likes99 downloads2y agoHugging Face29BioLaySumm /BioLaySumm2025-LaymanRRG-opensource-tracktext100K<n<1M0 likes78 downloads1y agoHugging Face30open-source-metrics /model-repos-stats Dataset Card for "model-repos-stats" More Information needed tabular100K<n<1M5 likes72 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.