datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tokenizers-dependents
tokenizers metrics
This dataset contains metrics about the huggingface/tokenizers package.
Number of repositories in the dataset: 11460
Number of packages in the dataset: 124
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 14 packages that have more than 1000 stars.
There are 41… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/tokenizers-dependents.transformers-dependents
transformers metrics
This dataset contains metrics about the huggingface/transformers package.
Number of repositories in the dataset: 27067
Number of packages in the dataset: 823
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 65 packages that have more than 1000 stars.
There are 140… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/transformers-dependents.Llasa_opensource_speech_data_160k_hours_tokenized
Update (2025-02-07): Our paper has been released!
This script is for merging tokenized speech datasets stored in memmap format. The input datasets can be combined to form larger training datasets.
import numpy as np
import os
def merge_memmap_datasets(dataset_dirs, output_dir):
# Ensure the output directory exists
os.makedirs(output_dir, exist_ok=True)
# Dataset splits to be merged
splits = ['train', 'val']
for split in splits:
shapes = []
seq_len =… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/Llasa_opensource_speech_data_160k_hours_tokenized.evaluate-dependents
evaluate metrics
This dataset contains metrics about the huggingface/evaluate package.
Number of repositories in the dataset: 106
Number of packages in the dataset: 3
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 1 packages that have more than 1000 stars.
There are 2 repositories… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/evaluate-dependents.gradio-dependents
Dataset Card for "gradio-dependents"
More Information needed
diffusers-dependents
diffusers metrics
This dataset contains metrics about the huggingface/diffusers package.
Number of repositories in the dataset: 160
Number of packages in the dataset: 2
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 0 packages that have more than 1000 stars.
There are 3 repositories… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/diffusers-dependents.optimum-dependents
optimum metrics
This dataset contains metrics about the huggingface/optimum package.
Number of repositories in the dataset: 19
Number of packages in the dataset: 6
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 0 packages that have more than 1000 stars.
There are 0 repositories that… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/optimum-dependents.accelerate-dependents
accelerate metrics
This dataset contains metrics about the huggingface/accelerate package.
Number of repositories in the dataset: 727
Number of packages in the dataset: 37
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 10 packages that have more than 1000 stars.
There are 16… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/accelerate-dependents.datasets-dependents
datasets metrics
This dataset contains metrics about the huggingface/datasets package.
Number of repositories in the dataset: 4997
Number of packages in the dataset: 215
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 22 packages that have more than 1000 stars.
There are 43… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/datasets-dependents.pip
Dataset Card for "pip"
More Information needed
wan22-animate-3k-opensource-data
Wan2.2 Animate Open Dataset Pack
This dataset repo stores the complete datasets/ directory used for the Wan2.2 TI2V 5B + One-to-All animate experiment.
The original tree contains more than 10,000 files in one directory, which Hugging Face git repositories reject as raw files. Therefore the dataset is stored as split tar shards.
Restore:
cat datasets.tar.part-* | tar -xf -
sha256sum -c SHA256SUMS
After extraction, the restored tree contains:… See the full description on the dataset page: https://huggingface.co/datasets/simbahuang/wan22-animate-3k-opensource-data.pytorch-image-models-dependents
pytorch-image-models metrics
This dataset contains metrics about the huggingface/pytorch-image-models package.
Number of repositories in the dataset: 3615
Number of packages in the dataset: 89
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 18 packages that have more than 1000… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/pytorch-image-models-dependents.rareburden-commons-open-source-snapshots
RareBurden Commons open source snapshots
This public preservation projection contains exact, hash-bound source files
whose observed terms affirmatively permit redistribution. Each source retains
its own licence; license: other is intentionally used because the collection
is not governed by one uniform licence.
Included:
Orphadata July 2026 alignment and epidemiology files — CC BY 4.0.
Exact MONDO release assets — CC BY 4.0. The currently receipt-bound history
covers v2026-08-04… See the full description on the dataset page: https://huggingface.co/datasets/edithatogo/rareburden-commons-open-source-snapshots.qmmit-open-source-agent-commit-index
Repository-Level Measurement of Self-Declared Coding-Agent Commit Signatures
Dataset release: 2026-09-18-v3.0Schema: 3.0.0
Release stamp: dataset 2026-09-18-v3.0 · ruleset sha256:b2e8889c66f72c18a61839f0bf1a9f77b5481ba2def044dba33c797b2f2bdcae · scanned 2026-09-16
Abstract
This dataset contains 2000 repository-level observations from public Git
repositories. Each observation estimates a lower bound on the proportion of
non-merge, non-infrastructure-bot commits… See the full description on the dataset page: https://huggingface.co/datasets/balrampandey/qmmit-open-source-agent-commit-index.color-classification-opensourcestarsissuesreinforcement-learning-checkpoint-downloadsai-opensource-2026
Ai Opensource 2026
Part of the LEGION Intelligence dataset collection.
Provider: LEGION Systems
Access: Requires approval — submit request below
Usage
from datasets import load_dataset
dataset = load_dataset("gemmozero/ai-opensource-2026")
API Access
Real-time access via LEGION API:
curl https://api.legion-api.com/incidents
API Docs · Pro Access €29/mo
License
CC BY-NC 4.0 — Research and non-commercial use only.
Commercial use… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-opensource-2026.opensource_mirrorfake_news_en_opensources
Dataset Card for "Fake News Opensources"
Dataset Description
Homepage: https://github.com/AndyTheFactory/FakeNewsDataset
Repository: https://github.com/AndyTheFactory/FakeNewsDataset
Point of Contact: Andrei Paraschiv
Dataset Summary
a consolidated and cleaned up version of the opensources Fake News dataset
Fake News Corpus comprises 8,529,090 individual articles, classified into 12 classes: reliable, unreliable, political, bias, fake, conspiracy… See the full description on the dataset page: https://huggingface.co/datasets/andyP/fake_news_en_opensources.os-world-modifiedopensource_100_TVG_caseopen_parallel_think_code_source
open_parallel_think_code_source
A large-scale code reasoning distillation dataset with 320,000 solution trajectories generated by 4 state-of-the-art thinking models across 10,000 unique coding problems.
Source / raw pool. This is the per-trajectory dataset. The packed parallel-thinking datasets derived from it are haowu89/open_parallel_think_code_full (full reasoning + solution) and haowu89/open_parallel_think_code_cot (solution only). Each trajectory's metadata carries… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/open_parallel_think_code_source.visual-question-answering-checkpoint-downloadshub-docs-dependents
Dataset Card for "hub-docs-dependents"
More Information needed
common-corpus-sample-open-sourcesafetensors-dependents
Dataset Card for "safetensors-dependents"
More Information needed
Evaluation_of_OpenSource_Models_for_PDF_Injection_Recognition
Injected PDFs - Model Evaluation
This repository holds the model evaluation stage of a project on detecting harmless-but-real
attack payloads injected into PDF files, together with the artefacts it produced for the
application.
Nothing is trained here. Seven off-the-shelf models are measured against the same 1,100 PDFs,
and the two winners are exported for the app to load.
Question
Candidates
Winner
Part A
Which files look like this one?
3 embedding models x 2 inputs… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/Evaluation_of_OpenSource_Models_for_PDF_Injection_Recognition.open-source-english-catalan-corpus
Dataset Card for open-source-english-catalan-corpus
Dataset Summary
Translation memory built from more than 180 open source projects. These include LibreOffice, Mozilla, KDE, GNOME, GIMP, Inkscape and many others. It can be used as translation memory or as training corpus for neural translators.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Catalan (ca)
English (en)
Dataset Structure
Data Instances
[More… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/open-source-english-catalan-corpus.
