datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Pytorch-Code-10K
Python/Pytorch Code Dataset
A collection of code from repos on The Stack, with captions generated by AI.
Dataset Description
This dataset contains Python code snippets sourced from open-source repositories that utilize PyTorch or Hugging Face Transformers. Each sample includes:
code: The raw Python source code (typically containing import torch, from torch import nn, or transformer-related imports)
caption: A natural language description generated by T5-Large… See the full description on the dataset page: https://huggingface.co/datasets/Monster-Code/Pytorch-Code-10K.IQA-PyTorch-Datasets
Description
This is the dataset repository used in the pyiqa toolbox. Please refer to Awesome Image Quality Assessment for details of each dataset
Example commandline script with huggingface-cli:
huggingface-cli download chaofengc/IQA-PyTorch-Datasets live.tgz --local-dir ./datasets --repo-type dataset
cd datasets
tar -xzvf live.tgz
Disclaimer for This Dataset Collection
This collection of datasets is compiled and maintained for academic, research, and educational… See the full description on the dataset page: https://huggingface.co/datasets/chaofengc/IQA-PyTorch-Datasets.alfa-scoring-trxhttps://ods.ai/competitions/dl-fintech-card-transactions
https://boosters.pro/championship/alfabattle2
tanuki-.nnue-pytorch-2024-07-30.1
Summary
Training Data for Shogi AI Development
Contents
tanuki-.nnue-pytorch-2024-07-30.1.7z.001...005 ... Training Data
The training data are provided in the YaneuraOu PackedSfenValue format.
This dataset was generated using tanuki-.nnue-pytorch-2024-07-30.1 with a search depth of 9.
The training data have not been shuffled. It is recommended to shuffle the training data before use. Additionally, positions within this dataset have not been replaced with the PV… See the full description on the dataset page: https://huggingface.co/datasets/nodchip/tanuki-.nnue-pytorch-2024-07-30.1.IQA-PyTorch-Datasets-metainfo
Description
This repo contains the meta information of datasets stored in chaofengc/IQA-PyTorch-Weights. They are used in the training codes of the pyiqa toolbox.
Disclaimer for Datasets Included
This collection of datasets is compiled and maintained for academic, research, and educational purposes. It is important to note the following points regarding the datasets included in this Collection:
Rights & Permissions: Each dataset in this Collection is the property of its… See the full description on the dataset page: https://huggingface.co/datasets/chaofengc/IQA-PyTorch-Datasets-metainfo.age-group-predictionhttps://ods.ai/competitions/sberbank-sirius-lesson
pytorch-nn-architectures-dataset
PyTorch Neural Network Architectures Dataset
608 PyTorch neural network implementations generated using GPT-5,
covering 7 architecture types, 4 task categories, 4 input data types,
and 4 complexity levels. All architectures are validated and ready to use.
What is in the dataset?
Each file is a standalone PyTorch class inheriting from torch.nn.Module, defining a complete and standalone neural network implementation. The prompt used to generate it is
included as… See the full description on the dataset page: https://huggingface.co/datasets/DNadia/pytorch-nn-architectures-dataset.retailhero-uplifthttps://ods.ai/competitions/x5-retailhero-uplift-modeling
pytorch-image-models-dependents
pytorch-image-models metrics
This dataset contains metrics about the huggingface/pytorch-image-models package.
Number of repositories in the dataset: 3615
Number of packages in the dataset: 89
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 18 packages that have more than 1000… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/pytorch-image-models-dependents.profiling-pytorch
Profiling in PyTorch Scripts
Holds all the scripts used for the series "Pofiling in PyTorch"
LocoTrack-pytorch-weightspytorch_rvt_rlbench_put_groceries_in_cupboard_bsz1transactions-genderhttps://www.kaggle.com/c/python-and-analyze-data-final-project/
pytorch-issues-dataset-cleanPyTorchConference2025_GithubRepos
PyTorch Conference 2025 GitHub Repos
I created a list of every GitHub repo mentioned during PyTorch Conference 2025 and Open Source AI Week.
pytorch_rvt_rlbench_open_drawer_bsz1gpu-cuda-pytorch-compatibility
PyTorch release, CUDA build and NVIDIA driver pairings
Canonical, always-current version: https://referencesource.org/gpu-cuda-pytorch-compatibility/
Machine-readable: https://referencesource.org/gpu-cuda-pytorch-compatibility/data.json — this mirror is a point-in-time copy.
Last verified: 2026-10-01
Stale after: 2026-12-28 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 45
One record per (PyTorch release, CUDA build)… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/gpu-cuda-pytorch-compatibility.rosbank-churnhttps://boosters.pro/championship/rosbank1/
github-pytorch-issues
Dataset Card for github-pytorch-issues
Dataset Summary
This dataset is a curated collection of GitHub issues from the PyTorch repository. Each entry includes the issue title, body, user, state, labels, comments, and other relevant fields that are useful for tasks such as text classification, semantic search, and question answering.
Supported Tasks and Leaderboards
The dataset supports the following tasks:
Open-domain Question Answering: Given a user query… See the full description on the dataset page: https://huggingface.co/datasets/mayankpuvvala/github-pytorch-issues.all-pytorch-codetanuki-.nnue-pytorch-2024-07-30.1_Shuffled_qsearch_by_haotanuki-.nnue-pytorch-2024-07-30.1をhaoでqsearch()を適応し、シャッフルしたもの。
alfa-scoring-bkihttps://ods.ai/competitions/dl-fintech-bki
pytorch-issues
Dataset Card for "pytorch-issues"
More Information needed
pytorch-issues-datasetpytorch-jetson-cn-mirror
版本名称下载地址说明
torch-2.3.0-cp310-cp310-linux_aarch64.whl下载JetPack 6.0 (L4T R36.2 / R36.3) + CUDA 12.2
torchvision-0.18.0a0+6043bc2-cp310-cp310-linux_aarch64.whl下载JetPack 6.0 (L4T R36.2 / R36.3) + CUDA 12.2
torchaudio-2.3.0+952ea74-cp310-cp310-linux_aarch64.whl下载JetPack 6.0 (L4T R36.2 / R36.3) + CUDA 12.2
torch-2.3.0-cp310-cp310-linux_aarch64.whl.zip下载JetPack 6.0 (L4T R36.2 / R36.3) + CUDA 12.4
torchvision-0.18.0a0+6043bc2-cp310-cp310-linux_aarch64.whl.zip下载JetPack 6.0 (L4T R36.2 / R36.3) +… See the full description on the dataset page: https://huggingface.co/datasets/futureflsl/pytorch-jetson-cn-mirror.pangu_pytorch
Auxiliary data (mask, constants, statistics...) and pytorch checkpoints required to reproduce performance of Pangu-Weather at 24-hour forecasting horizon.
pytorch-semantic-dataset-fixed
PyTorch Semantic Code Dataset
A semantically-enriched Python code dataset combining syntactic tokenization with deep semantic analysis from Language Server Protocol (LSP) tools.
🎯 Overview
This dataset enhances tokenized Python code with semantic embeddings derived from static analysis tools (Tree-sitter + Jedi), providing models with both syntactic and semantic understanding of code symbols. Each token in the code is aligned with rich semantic information including type… See the full description on the dataset page: https://huggingface.co/datasets/ant-des/pytorch-semantic-dataset-fixed.ttrshttps://arxiv.org/abs/2110.05589
flchain
Dataset Card for "flchain"
More Information needed
pytorch-precision-experiments
