datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PIDray
Dataset Card for pidray
PIDray is a large-scale dataset which covers various cases in real-world scenarios for prohibited item detection, especially for deliberately hidden items. The dataset contains 12 categories of prohibited items in 47, 677 X-ray images with high-quality annotated segmentation masks and bounding boxes.
This is a FiftyOne dataset with 9482 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/PIDray.digitize-pid-yolo
Digitize-PID (Symbols only), YOLO format
Note: I am not the author of this dataset
An annotated synthetic dataset, Dataset-P&ID, of 500 P&IDs with incorporates different
types of noise and complex symbols. This dataset contains only the symbols, i.e., under
the object detection task.
Paliwal, S., Jain, A., Sharma, M., & Vig, L. (2021). Digitize-PID: Automatic Digitization
of Piping and Instrumentation Diagrams. ArXiv, abs/2109.03794.
Original data repository:… See the full description on the dataset page: https://huggingface.co/datasets/hamzas/digitize-pid-yolo.pid_lines_dataset
P&ID Line Detection Dataset
This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams)
with line segment annotations for line detection and segmentation tasks.
Dataset Structure
Each sample contains:
file_name: Image filename
source_image_idx: Index of the original P&ID image
crop_idx: Index of this crop from the source image
width: Crop width in pixels
height: Crop height in pixels
lines: Dictionary with:
segments: List of line segments as [x1… See the full description on the dataset page: https://huggingface.co/datasets/prasatee/pid_lines_dataset.pid_lines_dataset
P&ID Line Detection Dataset
This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams)
with line segment annotations for line detection and segmentation tasks.
Dataset Structure
Each sample contains:
file_name: Image filename
source_image_idx: Index of the original P&ID image
crop_idx: Index of this crop from the source image
width: Crop width in pixels
height: Crop height in pixels
lines: Dictionary with:
segments: List of line… See the full description on the dataset page: https://huggingface.co/datasets/samiksha9874/pid_lines_dataset.metal_part_sort_v10_plus_extra_20260702
metal_part_sort_v10_plus_extra_20260702
Merged LeRobot v2.1 dataset for Unitree G1 + Inspire DFX metal part sorting.
Base dataset: PID0930/metal_part_sort_v10 @ 17fe43673ca30e840b1baa6e05c1f34b7dd43b91
Extra dataset: PID0930/metal_part_sort_v10_extra_20260701 @ 09c5a9113830443f99a8d91c11fd5754b238fc6a
Episodes: 329
Frames: 110966
Cameras: external, left wrist, right wrist, plus left_high/head slot retained for compatibility
GR00T training uses external as the logical head… See the full description on the dataset page: https://huggingface.co/datasets/PID0930/metal_part_sort_v10_plus_extra_20260702.pi-dev-plugins
Pi Coding Agent Plugins
Dataset contains metadata for approximately 5500 Pi Coding Agent plugins gathered from https://pi.dev/packages on 2026-10-05. Row example:
{
"name": "pi-mcp-adapter",
"description": "MCP (Model Context Protocol) adapter extension for Pi coding agent",
"types": [
"extension"
],
"author": "nicopreme",
"downloads_monthly": 1507084,
"downloads_label": "1.5M/mo",
"published_label": "4h ago",
"published_ms": 1791251148391,
"detail_url":… See the full description on the dataset page: https://huggingface.co/datasets/kth8/pi-dev-plugins.PIDdatasetEgoManipulation-RGBD-Beta
EgoManipulation RGB-D Beta
EgoManipulation RGB-D Beta is a collection of egocentric manipulation recordings containing RGB and depth streams time-synchronized using timestamps from a shared device clock.
A Parquet catalog provides recording-level metadata and previews for browsing the dataset without downloading the full MCAP collection.
Data recordings are stored as MCAP files, with one selected recording chunk per file.
Repository layout
README.md
data/… See the full description on the dataset page: https://huggingface.co/datasets/PI-Dojo/EgoManipulation-RGBD-Beta.metal_part_sort_v2stocks-PIDILITIND-1D-candlespi-diff-review
Coding agent session traces for badlogicgames/pi-diff-review
This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line… See the full description on the dataset page: https://huggingface.co/datasets/badlogicgames/pi-diff-review.pid-icons-mergednigerian-pidgin-1.0
Language:
- Nigerian Pidgin English (West African Pidgin variant)
Dataset Description
Dataset Summary
The Nigerian Pidgin ASR dataset (v1.0) is the first publicly released speech-to-text corpus for Nigerian Pidgin English, a widely spoken lingua franca across Nigeria and West Africa. This dataset comprises over 3,000 audio recordings paired with sentence-level transcriptions, recorded by native speakers across different genders and age groups. It is tailored for… See the full description on the dataset page: https://huggingface.co/datasets/asr-nigerian-pidgin/nigerian-pidgin-1.0.cloth_folding_v2_pidpjqBench
jqBench
Benchmarks for generating executable JSON queries and transformations from natural-language instructions, examples, or schemas.
Configurations
Configuration
Rows
Description
jqStack
1,496
Generated and reviewed Stack Overflow tasks
jqStackEasy
12,988
Additional collected tasks excluding reviewed benchmark identifiers
jqSpider
859
Spider-derived tasks with shared JSON databases and schemas
All configurations use the split name data.… See the full description on the dataset page: https://huggingface.co/datasets/PidgeyUsedGust/jqBench.pi-diff-review
Coding agent session traces for badlogicgames/pi-diff-review
This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each sessions/*.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/pi-diff-review.sample_pidp9jalingo-reviewed-pidginpad-pid-geospatial
Geospatial data use in World Bank PADs and PIDs
The data behind How often do World Bank projects use geospatial data?.
Snapshot 938866a81eab, annotations to 2026-09-25 21:57 UTC. The page shows the same snapshot id.
Headline
Out of: those that use data
Out of: all
Projects
1,165 of 1,932 (60.3%)
1,165 of 2,261 (51.5%)
Documents
1,637 of 3,544 (46.2%)
1,637 of 4,494 (36.4%)
A project uses geospatial data when at least one data mention in the… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/pad-pid-geospatial.pi_double_cam_test_20261001_145138This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/oscarcole03/pi_double_cam_test_20261001_145138.pidu-data
PIDU simulated datasets
Simulated infrared spectra (Transfer Matrix Method, random Drude-Lorentz layer and oscillator
parameters) with their ground-truth parameters, used to train and test the models of
Physics-infused Deep Unfolding for Automated Material Parameter Extraction of Optical Data
(Koumans, Stevens, van Sloun, van Mechelen).
Code: https://github.com/MKoumans/PIDU · Checkpoints: MKoumans/pidu-optical-parameter-extraction
Contents
Path
Content… See the full description on the dataset page: https://huggingface.co/datasets/MKoumans/pidu-data.digitize-pid-yolo
Digitize-PID (Symbols only), YOLO format
Note: I am not the author of this dataset
An annotated synthetic dataset, Dataset-P&ID, of 500 P&IDs with incorporates different
types of noise and complex symbols. This dataset contains only the symbols, i.e., under
the object detection task.
Paliwal, S., Jain, A., Sharma, M., & Vig, L. (2021). Digitize-PID: Automatic Digitization
of Piping and Instrumentation Diagrams. ArXiv, abs/2109.03794.
Original data repository:… See the full description on the dataset page: https://huggingface.co/datasets/SherifAhmed/digitize-pid-yolo.PID2GraphPidginMed-ASR
PidginMed-ASR
PidginMed-ASR is a Nigerian Pidgin English medical speech benchmark for automatic speech recognition.
Dataset Structure
The dataset contains audio recordings paired with Nigerian Pidgin transcriptions and English source translations.
Column
Description
audio
Path to audio file
transcription
Nigerian Pidgin transcript
english_translation
Original English text
speaker_id
Anonymized speaker ID
duration
Audio duration
split
Train, validation… See the full description on the dataset page: https://huggingface.co/datasets/AnnonSubmit/PidginMed-ASR.pid-object-detection
Dataset Labels
['ball-valve', 'butterfly-valve', 'centrifugal-pump', 'check-valve', 'gate-valve']
Number of Images
{'valid': 12, 'test': 12, 'train': 128}
How to Use
Install datasets:
pip install datasets
Load the dataset:
from datasets import load_dataset
ds = load_dataset("DanielCerda/pid-object-detection", name="full")
example = ds['train'][0]
Roboflow Dataset Page… See the full description on the dataset page: https://huggingface.co/datasets/DanielCerda/pid-object-detection.T-PID
In the name of Allah
Sample of Toorintan-Persian Informal Dataset (T-PID)
Sample data is available in the train, dev, and test folders. Corresponding metadata can be found in the train.tsv, dev.tsv, and test.tsv files, which are available for download.
The full dataset is currently under review for publication as part of an academic paper. Once accepted, we will release the complete dataset for research purposes.
📩 To request access to the full dataset in the meantime… See the full description on the dataset page: https://huggingface.co/datasets/Toorintan/T-PID.metal_part_sort_v6
metal_part_sort_v6
LeRobot v2.1 dataset for the Unitree G1 + Inspire DFX metal part sorting task.
Episodes: 39
Frames: 32,049
FPS: 30
Raw snapshot: /home/y-kosugi/datasets/metal_part_sort_v6_all_raw_20260624/metal_part_sort_v6
Camera streams are stored as:
observation.images.cam_left_high: robot head camera, recorded for audit only
observation.images.cam_external: external camera
observation.images.cam_left_wrist: left wrist camera
observation.images.cam_right_wrist: right… See the full description on the dataset page: https://huggingface.co/datasets/PID0930/metal_part_sort_v6.metal_part_sort_v3
metal_part_sort_v3
LeRobot v2.1 dataset for Unitree G1 + Inspire DFX metal part sorting.
Episodes: 95
Frames: 46748
Cameras: head, left wrist, right wrist
State/action: 26D, with hand dimensions 14:26
Hand-state convention: observation.state[t, 14:26] = action[t-1, 14:26] for t > 0; frame 0 uses same-frame hand action.
Failed recordings moved to trash were excluded before conversion.
Pidgin_ASR_Dataset_Combined
Naija-ASR-Corpus v1.0 (NAC-v1.0)
A Foundational Automatic Speech Recognition Corpus for Nigerian Pidgin (Naija, PCM)
📌 Dataset Summary
Naija-ASR-Corpus (NAC-v1.0) is a speech dataset derived from the Universal Dependencies Naija Spoken Corpus (UD_Naija-NSC).
The NAC Team processed the original long-form recordings by:
Segmenting the audio into sentence-level clips.
Transcribing/Aligning the text to create paired audio-text data suitable for ASR training.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/timniel/Pidgin_ASR_Dataset_Combined.g1-inspire-pick-cube-gr00t-lerobot
Unitree G1 Inspire Broccoli Plush To Red Tray
GR00T LeRobot v2 dataset converted from Unitree-style JSON episodes and JPEG frames.
Robot: Unitree G1 with Inspire DFX56 hands
Task: Pick up the broccoli plush with the right hand and place it on the red tray.
Episodes: 9
Frames: 3979
Format: GR00T-flavored LeRobot v2 with meta/modality.json
Use with Isaac-GR00T N1.7 via --embodiment-tag NEW_EMBODIMENT and
examples/G1Inspire/g1_inspire_config.py.
