datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
energy-consumption-hourly-spaincc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13vital
Vital PresetShare Renders
Rendered Vital presets scraped from PresetShare.
sample rate: 22050
render duration: 6.0s
MIDI note: 72 (C4)
note duration: 5.0s
velocity: 100
Files are organized under by_type/<sound-type>/<preset-id>_<name>/ with:
preset.vital
preview.mp3
vital-render.wav
metadata.json
See manifest.jsonl and summary.json for run metadata.
sam2-vit-bflickr30k_clip-ViT-B-32-caption_pairssam2-vit-senergy-consumption-weather-hourly-spaincls_dinov3-vith16plus_in22ksarscov2_vitro_touret
Dataset Details
Dataset Description
An in-vitro screen of the Prestwick chemical library composed of 1,480
approved drugs in an infected cell-based assay.
Curated by:
License: CC BY 4.0
Dataset Sources
corresponding publication
Data source
Citation
BibTeX:
@article{Touret2020,
doi = {10.1038/s41598-020-70143-6},
url = {https://doi.org/10.1038/s41598-020-70143-6},
year = {2020},
month = aug,
publisher = {Springer Science and Business Media LLC}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/sarscov2_vitro_touret.cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15openaccess-embeddings-openclip-vitg14
metmuseum/openaccess-embeddings-openclip-vitg14
Image embeddings for every public-domain artwork in metmuseum/openaccess, produced by laion/CLIP-ViT-bigG-14-laion2B-s39B-b160K.
Column
Type
Notes
objectID
int64
Primary key — matches objectID in metmuseum/openaccess
embedding
list<float32>
L2-normalised, dim = 1280
model
string
Source model id
dim
int32
Embedding dimension
Image bytes are not stored here; join against the main dataset to recover them.Embedding… See the full description on the dataset page: https://huggingface.co/datasets/metmuseum/openaccess-embeddings-openclip-vitg14.structured-vitalsmscoco_train_2014_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15adaptive_rag_2wikimultihopqamscoco_train_2014_openai_clip-vit-base-patch32_image_caption_retrieval_pairs_2022-09-01osworld_tasks_fileslaion-1m-vit-h-14gigahands-vitra-mano
GigaHands → VITRA Stage-1, official-MANO annotations
VITRA Stage-1 hand annotations for GigaHands, with all joint positions taken from GigaHands'
official MANO fit instead of mixing in triangulated keypoints. Annotations only — no videos
(get those from GigaHands; the mapping is described in §5).
episodes
13,247 (train 11,904 / test 1,343)
frames
3,395,733
camera
brics-odroid-001_cam0 (static rig; one constant extrinsic per scene)
source
GigaHands params/ +… See the full description on the dataset page: https://huggingface.co/datasets/yuzhench/gigahands-vitra-mano.heloc
HELOC (Home Equity Line of Credit)
The HELOC dataset from FICO.
Each entry in the dataset is a line of credit, typically offered by a bank as a percentage of home equity (the difference between the current market value of a home and its purchase price).
The customers in this dataset have requested a credit line in the range of $5,000 - $150,000.
The fundamental task is to use the information about the applicant in their credit report to predict whether they will repay their HELOC… See the full description on the dataset page: https://huggingface.co/datasets/vitaliykinakh/heloc.yam-vital-left-hand
vital_left_arm
Teleoperation dataset: Grasp the vital, and insert it into the stand.
Collected with limb on yam arms.
Dataset summary
Field
Value
Robot
yam
Episodes
350
Total frames
113,165
FPS
30 Hz
Task
Grasp the vital, and insert it into the stand.
Format
LeRobot v3.0
Cameras
Name
head_camera
left_wrist_camera
Video codec: hevc.
State space (observation.state, shape [7])
Index
Name… See the full description on the dataset page: https://huggingface.co/datasets/Sichang0621/yam-vital-left-hand.bridgev2-vita-toykitchen-manifests
BridgeV2 VITA ToyKitchen-like Manifests
This repository contains manifest files for a reconstructed VITA-style BridgeV2 ToyKitchen-like pick-and-place subset.
Source dataset
The source dataset is:
Gaugou/BridgeV2
This repository does not duplicate the original BridgeV2 videos. It provides episode IDs and metadata for selecting the subset from the source dataset.
Split
Train: 2,986 episodes
Test: 287 episodes
Total selected: 3,273 episodes
Selection… See the full description on the dataset page: https://huggingface.co/datasets/praedico/bridgev2-vita-toykitchen-manifests.synthetic-fraud-detectionr3al-vit-quantization-codex-trace
R3AL ViT Quantization — Codex Agent Trace
Codex session trace for installing the R3AL CLI and agent skill, exporting
google/vit-base-patch16-224 to ONNX, performing dynamic INT8 post-training
quantization on R3AL, and evaluating model size, Apple-arm64 CPU latency, and
prediction fidelity on a 100-image ImageNet validation sample.
The original Codex JSONL format is preserved for Hugging Face's native Agent
Trace viewer. Credential values, email addresses, unrelated Gmail/Slack… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/r3al-vit-quantization-codex-trace.eval_vit_rectified_seenpos_8degThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 15,
"total_frames": 2433,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/miladgholami/eval_vit_rectified_seenpos_8deg.pnp_vita_21_pi05This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 21,
"total_frames": 6213,
"total_tasks": 1,
"total_videos": 42,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lancel2025/pnp_vita_21_pi05.SO101_pillbox_vitaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 5,
"total_frames": 1422,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/praedico/SO101_pillbox_vita.mscoco_train_2014_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13eval_vit_rectified_seenpos_22degThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 15,
"total_frames": 2596,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/miladgholami/eval_vit_rectified_seenpos_22deg.flickr30k_CLIP_ViT-B-32_subset_pairs_SimCSE_similarity_copycs280-synthetic-circles-v2-embeddings-vjepa21-vitg-384
