datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tcga-wsi-uni2h-features
TCGA WSI UNI2H Features
Dataset Summary
This dataset provides tile-level UNI2-h embeddings extracted from TCGA whole-slide images (WSIs) using a reproducible, auditable pipeline designed for computational pathology research.
Data is organized by project (for example TCGA-HNSC) and currently exposes:
features/ containing H5 feature files with tile-level embeddings
vis/ containing overlay images for quality inspection and pipeline verification
[!IMPORTANT]
Unlike the… See the full description on the dataset page: https://huggingface.co/datasets/W8Yi/tcga-wsi-uni2h-features.openwakeword_featuresThis dataset contains precomputed audio features designed for use with the openWakeWord library.
Specifically, they are intended to be used as general purpose negative data (that is, data that does not contain the target wake word/phrase) for training custom openWakeWord models.
The individual .npy files in this dataset are not original audio data, but rather are low dimensional audio features produced by a pre-trained speech embedding model from Google.
openWakeWord uses these features as… See the full description on the dataset page: https://huggingface.co/datasets/davidscripka/openwakeword_features.SN-Features
SoccerNet Features
Pre-extracted per-game features for the SoccerNet benchmark, structured as <league>/<season>/<game>/<file>, one file per game half (1_.../2_...).
This main branch holds no data — each feature type lives on its own branch so you only download what you need:
Branch
Files
Description
baidu-soccer-embeddings
{1,2}_baidu_soccer_embeddings.npy
Frame embeddings from baidu-research/vidpress-sports, used by the Action Spotting and Dense Video Captioning 2023… See the full description on the dataset page: https://huggingface.co/datasets/SoccerNet/SN-Features.conch_v15_features
CONCH v1.5 Patch Features for TCGA and CPTAC
Pre-extracted patch-level embeddings from the CONCH v1.5 pathology foundation model for 11,760 whole-slide images (WSIs): 9,838 from TCGA (32 projects) and 1,922 from CPTAC (9 cohorts).
Features were extracted with TRIDENT. They are meant for weakly supervised slide-level tasks, such as multiple-instance learning (MIL) for classification, survival or biomarker prediction, without having to download or process the raw WSIs.
These… See the full description on the dataset page: https://huggingface.co/datasets/sofieneb/conch_v15_features.bulk-cc12m-features
bulk-cc12m-features — ten teacher towers over CC12M, plus their consensus
Precomputed image-tower features for 10,968,539 CC12M images (all 2,176
shards of
pixparse/cc12m-wds)
from ten independent teacher extractions — eight CLIP variants across
three pretraining corpora and two model scales, SigLIP, and DINOv3 — plus
one derived consensus target.
About 110 million feature vectors, roughly 130 GPU-hours of extraction,
so that a student can be distilled against any of these… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/bulk-cc12m-features.livekit_wakeword_featuresThis dataset contains precomputed audio features designed for use with the openWakeWord library.
Specifically, they are intended to be used as general purpose negative data (that is, data that does not contain the target wake word/phrase) for training custom openWakeWord models.
The individual .npy files in this dataset are not original audio data, but rather are low dimensional audio features produced by a pre-trained speech embedding model from Google.
openWakeWord uses these features as… See the full description on the dataset page: https://huggingface.co/datasets/binhpham/livekit_wakeword_features.ucf-crime-clip-featuresucf-crime-processed-featuresfeatures-dinov3-vith16plus-224-imagenet-22k-wdsframeflow-ltx-resfix-r4-features-12h-20261002
resfix_r4_features_12h
Joint protein representation autoencoder ablation. See config.yaml, provenance.json and status.json for the exact configuration and progress. W&B: https://wandb.ai/gaorory-ucla-team/frameflow-ltx/runs/8cfxi4uo
Exact FP32 model weights are retained every 5 optimizer steps inside immutable checkpoints/.tar shards. Each member is a torch checkpoint containing vae, config, step, and train_seconds. checkpoints.jsonl records member names and SHA-256 hashes.… See the full description on the dataset page: https://huggingface.co/datasets/raftbioworks/frameflow-ltx-resfix-r4-features-12h-20261002.embodied_features_and_demos_liberoDataset for Embodied Chain-of-Thought Reasoning for LIBERO-90, as used by ECoT-Lite.
TFDS Demonstration Data
The TFDS dataset contains successful demonstration trajectories for LIBERO-90 (50 trajectories for each of 90 tasks). It was created by rolling out the actions provided in the original LIBERO release and filtering out all unsuccessful ones, leaving 3917 successful demo trajectories. This is done via a modified version of a script from the MiniVLA codebase. In addition to… See the full description on the dataset page: https://huggingface.co/datasets/Embodied-CoT/embodied_features_and_demos_libero.Mintaka_Graph_Features_T5-xl-ssm
Dataset Card for "Mintaka_Graph_Features_T5-xl-ssm"
More Information needed
UNI2-h-features
Dataset Card for UNI2-h Pretrained Features
This dataset card provides the UNI2-h features for TCGA, CPTAC, and PANDA datasets with patch size 256 x 256 pixels at 20x magnification.
Requesting Access
As mentioned in the gated prompt, you must agree to the outlined terms of use, with the primary email for your HuggingFace account matching your institutional email. If your primary email is a personal email (@gmail/@hotmail/@qq) your request will be denied. To fix this, you… See the full description on the dataset page: https://huggingface.co/datasets/MahmoodLab/UNI2-h-features.features_mavos_completeztf-dr3-m31-featuresfeatures_dict_x40_subset_v3precomputed_audio_featuresAV-CIL_featuresbehaviour1k-Qwen3-features
BEHAVIOR-1K Qwen3 skill features
Per-frame conditioned features e_t = Phi(f_t, L_sub^(j), L), mean-pooled primitive skill latents S_j, aligned proprioception q_t, actions a_t, and subtask progress p_t.
These are the inputs and targets for a Primitive Skill Composer VLA Skill Predictor.
Ground-truth primitives come from BEHAVIOR-1K's hand-authored
primitive_annotation, so the segmentation is human-labelled rather than
predicted, and nothing here depends on a keyframe detector.… See the full description on the dataset page: https://huggingface.co/datasets/erl-hub/behaviour1k-Qwen3-features.TissueMNIST-224-full-gpt5nano-with-vlm-features
TissueMNIST 224 Full Train Val with GPT-5-nano VLM Features
The full TissueMNIST train and validation splits with categorical morphology features generated by GPT-5-nano. Test is included as the full TissueMNIST passthrough split with null vlm_model_name and placeholder vlm_feature values for schema consistency.
This dataset is derived from the official MedMNIST TissueMNIST 224px data.
The VLM feature labels are categorical privileged-information annotations
for CS231N VLM-LUPI… See the full description on the dataset page: https://huggingface.co/datasets/Koa-Chang/TissueMNIST-224-full-gpt5nano-with-vlm-features.vitra-dinotxt-featuresbulk-coco-featuresHere exists the bulk prepared sets for coco 2017.
With this I will begin testing the first WIDE ViT-Beatrix, ViT-Zana, ViT-Beatrix-DualStream, Clip-Vit-Beatrix, GeoVit-Beans and more.
These wide vits will be using new forms of formula meant to fuse structural behaviors together which exist on multiple different manifolds simultaneously.
These upcoming experiments will be with established SOTA-based processes adopted and modulated for geofractal behavior from multiple transfer learning… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/bulk-coco-features.360x_dataset_featuresAlgonautsDS-features
Saved Features for Algonauts '25 Dataset
This repository contains pre-extracted features for the Algonauts Challenge dataset using baseline models.
Features Overview
The developer_kit directory contains features extracted for the entire dataset using the following models:
Video Features
Model: SlowFast R50
Extracts spatiotemporal features from video frames
Captures motion and appearance information
Audio Features
Model: MFCC (Mel-frequency… See the full description on the dataset page: https://huggingface.co/datasets/medarc/AlgonautsDS-features.THGS-lerf-ovs-language-features
THGS — LERF-OVS language_features (precomputed)
Precomputed per-view language features for the 4 LERF-OVS scenes
(figurines, ramen, teatime, waldo_kitchen), used as input to the
THGS pipeline (merge_proj.py, Stage 3 replay).
For each training image there are two files:
file
content
frame_XXXXX_s.npy
per-view SAM segmentation maps (4-level, LangSplat-variant SAM)
frame_XXXXX_f.npy
per-mask CLIP features
Generated with scripts/image_encoding.py using the… See the full description on the dataset page: https://huggingface.co/datasets/JUNHAKBAE/THGS-lerf-ovs-language-features.morph_features
UniMorph + UniSegments Morph Data
This dataset pairs UniMorph inflectional features with UniSegments segmentations. For languages without UniSegments coverage, segmentation defaults to the unsegmented word form itself.
This resource is a necessary component for evaluating Tokenizer Morphological Plausibility, as introduced in Tokenizer Morphological Plausibility (https://arxiv.org/abs/2601.18536). The data generation process follows the implementation provided in the official… See the full description on the dataset page: https://huggingface.co/datasets/SHENJJ1017/morph_features.autodata-climbmix-features
AutoData — ClimbMix Feature Bank
Per-document annotations for the ClimbMix pre-training pool, used by the
data-selection recipes in WecoAI/AutoData.
Two banks are released:
Directory
Coverage
Contents
Size
/ (root)
full pool — 553,155,584 docs
lexical + perplexity + topic/format
~22.7 GiB
reasoning_53shards/
first 4,485,120 docs
Gemini reasoning/error annotations
~26 MiB
License and provenance
This dataset contains derived per-document… See the full description on the dataset page: https://huggingface.co/datasets/WecoAI/autodata-climbmix-features.malicious-website-features-2.4MImportant Notice:
A subset of the URL dataset is from Kaggle, and the Kaggle datasets contained 10%-15% mislabelled data. See this dicussion I opened for some false positives. I have contacted Kaggle regarding their erroneous "Usability" score calculation for these unreliable datasets.
The feature extraction methods shown here are not robust at all in 2023, and there're even silly mistakes in 3 functions: not_indexed_by_google, domain_registration_length, and age_of_domain.
The features… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/malicious-website-features-2.4M.librispeech-phoneme-featuresspotify_audio_features
Spotify Tracks & Audio Features Dataset
Overview
This dataset contains a comprehensive collection of Spotify tracks, combining rich audio feature analysis with track metadata. It is formatted as a high-performance Parquet dataset (ZStandard compressed), optimized for large-scale tabular analysis, machine learning, and recommender system research.
Data Source
The raw data for this dataset was originally gathered and hosted by Anna's Archive.
Original Blog Post:… See the full description on the dataset page: https://huggingface.co/datasets/ozefe/spotify_audio_features.
