Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01W8Yi /tcga-wsi-uni2h-features TCGA WSI UNI2H Features Dataset Summary This dataset provides tile-level UNI2-h embeddings extracted from TCGA whole-slide images (WSIs) using a reproducible, auditable pipeline designed for computational pathology research. Data is organized by project (for example TCGA-HNSC) and currently exposes: features/ containing H5 feature files with tile-level embeddings vis/ containing overlay images for quality inspection and pipeline verification [!IMPORTANT] Unlike the… See the full description on the dataset page: https://huggingface.co/datasets/W8Yi/tcga-wsi-uni2h-features.imageimage-feature-extractionn<1K13 likes19k downloads7mo agoHugging Face02davidscripka /openwakeword_featuresThis dataset contains precomputed audio features designed for use with the openWakeWord library. Specifically, they are intended to be used as general purpose negative data (that is, data that does not contain the target wake word/phrase) for training custom openWakeWord models. The individual .npy files in this dataset are not original audio data, but rather are low dimensional audio features produced by a pre-trained speech embedding model from Google. openWakeWord uses these features as… See the full description on the dataset page: https://huggingface.co/datasets/davidscripka/openwakeword_features.2 likes17k downloads3y agoHugging Face03SoccerNet /SN-Features SoccerNet Features Pre-extracted per-game features for the SoccerNet benchmark, structured as <league>/<season>/<game>/<file>, one file per game half (1_.../2_...). This main branch holds no data — each feature type lives on its own branch so you only download what you need: Branch Files Description baidu-soccer-embeddings {1,2}_baidu_soccer_embeddings.npy Frame embeddings from baidu-research/vidpress-sports, used by the Action Spotting and Dense Video Captioning 2023… See the full description on the dataset page: https://huggingface.co/datasets/SoccerNet/SN-Features.other0 likes4.3k downloads1mo agoHugging Face04sofieneb /conch_v15_featuresgated CONCH v1.5 Patch Features for TCGA and CPTAC Pre-extracted patch-level embeddings from the CONCH v1.5 pathology foundation model for 11,760 whole-slide images (WSIs): 9,838 from TCGA (32 projects) and 1,922 from CPTAC (9 cohorts). Features were extracted with TRIDENT. They are meant for weakly supervised slide-level tasks, such as multiple-instance learning (MIL) for classification, survival or biomarker prediction, without having to download or process the raw WSIs. These… See the full description on the dataset page: https://huggingface.co/datasets/sofieneb/conch_v15_features.10K<n<100K0 likes3.7k downloads8d agoHugging Face05AbstractPhil /bulk-cc12m-features bulk-cc12m-features — ten teacher towers over CC12M, plus their consensus Precomputed image-tower features for 10,968,539 CC12M images (all 2,176 shards of pixparse/cc12m-wds) from ten independent teacher extractions — eight CLIP variants across three pretraining corpora and two model scales, SigLIP, and DINOv3 — plus one derived consensus target. About 110 million feature vectors, roughly 130 GPU-hours of extraction, so that a student can be distilled against any of these… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/bulk-cc12m-features.text100M<n<1B0 likes2.9k downloads2mo agoHugging Face06binhpham /livekit_wakeword_featuresThis dataset contains precomputed audio features designed for use with the openWakeWord library. Specifically, they are intended to be used as general purpose negative data (that is, data that does not contain the target wake word/phrase) for training custom openWakeWord models. The individual .npy files in this dataset are not original audio data, but rather are low dimensional audio features produced by a pre-trained speech embedding model from Google. openWakeWord uses these features as… See the full description on the dataset page: https://huggingface.co/datasets/binhpham/livekit_wakeword_features.0 likes2.6k downloads7mo agoHugging Face07myzhao1999 /ucf-crime-clip-features0 likes2.6k downloads2y agoHugging Face08Kavindu1124 /ucf-crime-processed-features0 likes2.2k downloads1mo agoHugging Face09Emanresu /features-dinov3-vith16plus-224-imagenet-22k-wdstext1M<n<10M0 likes2.1k downloads11mo agoHugging Face10raftbioworks /frameflow-ltx-resfix-r4-features-12h-20261002 resfix_r4_features_12h Joint protein representation autoencoder ablation. See config.yaml, provenance.json and status.json for the exact configuration and progress. W&B: https://wandb.ai/gaorory-ucla-team/frameflow-ltx/runs/8cfxi4uo Exact FP32 model weights are retained every 5 optimizer steps inside immutable checkpoints/.tar shards. Each member is a torch checkpoint containing vae, config, step, and train_seconds. checkpoints.jsonl records member names and SHA-256 hashes.… See the full description on the dataset page: https://huggingface.co/datasets/raftbioworks/frameflow-ltx-resfix-r4-features-12h-20261002.0 likes2k downloads6d agoHugging Face11Embodied-CoT /embodied_features_and_demos_liberoDataset for Embodied Chain-of-Thought Reasoning for LIBERO-90, as used by ECoT-Lite. TFDS Demonstration Data The TFDS dataset contains successful demonstration trajectories for LIBERO-90 (50 trajectories for each of 90 tasks). It was created by rolling out the actions provided in the original LIBERO release and filtering out all unsuccessful ones, leaving 3917 successful demo trajectories. This is done via a modified version of a script from the MiniVLA codebase. In addition to… See the full description on the dataset page: https://huggingface.co/datasets/Embodied-CoT/embodied_features_and_demos_libero.robotics4 likes1.9k downloads6mo agoHugging Face12s-nlp /Mintaka_Graph_Features_T5-xl-ssm Dataset Card for "Mintaka_Graph_Features_T5-xl-ssm" More Information needed tabular100K<n<1M0 likes1.9k downloads3y agoHugging Face13MahmoodLab /UNI2-h-featuresgated Dataset Card for UNI2-h Pretrained Features This dataset card provides the UNI2-h features for TCGA, CPTAC, and PANDA datasets with patch size 256 x 256 pixels at 20x magnification. Requesting Access As mentioned in the gated prompt, you must agree to the outlined terms of use, with the primary email for your HuggingFace account matching your institutional email. If your primary email is a personal email (@gmail/@hotmail/@qq) your request will be denied. To fix this, you… See the full description on the dataset page: https://huggingface.co/datasets/MahmoodLab/UNI2-h-features.42 likes1.8k downloads2y agoHugging Face14acroitoru /features_mavos_complete0 likes1.3k downloads1y agoHugging Face15snad-space /ztf-dr3-m31-featurestabular10K<n<100K0 likes1.3k downloads2y agoHugging Face16aaljuhani /features_dict_x40_subset_v30 likes1k downloads7mo agoHugging Face17DavidErikMollberg /precomputed_audio_features0 likes1k downloads1y agoHugging Face18wiberg /AV-CIL_features0 likes979 downloads3y agoHugging Face19erl-hub /behaviour1k-Qwen3-features BEHAVIOR-1K Qwen3 skill features Per-frame conditioned features e_t = Phi(f_t, L_sub^(j), L), mean-pooled primitive skill latents S_j, aligned proprioception q_t, actions a_t, and subtask progress p_t. These are the inputs and targets for a Primitive Skill Composer VLA Skill Predictor. Ground-truth primitives come from BEHAVIOR-1K's hand-authored primitive_annotation, so the segmentation is human-labelled rather than predicted, and nothing here depends on a keyframe detector.… See the full description on the dataset page: https://huggingface.co/datasets/erl-hub/behaviour1k-Qwen3-features.tabularrobotics1K<n<10K0 likes907 downloads2mo agoHugging Face20Koa-Chang /TissueMNIST-224-full-gpt5nano-with-vlm-features TissueMNIST 224 Full Train Val with GPT-5-nano VLM Features The full TissueMNIST train and validation splits with categorical morphology features generated by GPT-5-nano. Test is included as the full TissueMNIST passthrough split with null vlm_model_name and placeholder vlm_feature values for schema consistency. This dataset is derived from the official MedMNIST TissueMNIST 224px data. The VLM feature labels are categorical privileged-information annotations for CS231N VLM-LUPI… See the full description on the dataset page: https://huggingface.co/datasets/Koa-Chang/TissueMNIST-224-full-gpt5nano-with-vlm-features.text100K<n<1M0 likes866 downloads4mo agoHugging Face21sjmathy /vitra-dinotxt-features0 likes789 downloads3mo agoHugging Face22AbstractPhil /bulk-coco-featuresHere exists the bulk prepared sets for coco 2017. With this I will begin testing the first WIDE ViT-Beatrix, ViT-Zana, ViT-Beatrix-DualStream, Clip-Vit-Beatrix, GeoVit-Beans and more. These wide vits will be using new forms of formula meant to fuse structural behaviors together which exist on multiple different manifolds simultaneously. These upcoming experiments will be with established SOTA-based processes adopted and modulated for geofractal behavior from multiple transfer learning… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/bulk-coco-features.timeseriesfeature-extraction1M<n<10M0 likes720 downloads10mo agoHugging Face23quchenyuan /360x_dataset_features0 likes673 downloads2y agoHugging Face24medarc /AlgonautsDS-features Saved Features for Algonauts '25 Dataset This repository contains pre-extracted features for the Algonauts Challenge dataset using baseline models. Features Overview The developer_kit directory contains features extracted for the entire dataset using the following models: Video Features Model: SlowFast R50 Extracts spatiotemporal features from video frames Captures motion and appearance information Audio Features Model: MFCC (Mel-frequency… See the full description on the dataset page: https://huggingface.co/datasets/medarc/AlgonautsDS-features.text10K<n<100K2 likes669 downloads1y agoHugging Face25JUNHAKBAE /THGS-lerf-ovs-language-features THGS — LERF-OVS language_features (precomputed) Precomputed per-view language features for the 4 LERF-OVS scenes (figurines, ramen, teatime, waldo_kitchen), used as input to the THGS pipeline (merge_proj.py, Stage 3 replay). For each training image there are two files: file content frame_XXXXX_s.npy per-view SAM segmentation maps (4-level, LangSplat-variant SAM) frame_XXXXX_f.npy per-mask CLIP features Generated with scripts/image_encoding.py using the… See the full description on the dataset page: https://huggingface.co/datasets/JUNHAKBAE/THGS-lerf-ovs-language-features.0 likes641 downloads4mo agoHugging Face26SHENJJ1017 /morph_features UniMorph + UniSegments Morph Data This dataset pairs UniMorph inflectional features with UniSegments segmentations. For languages without UniSegments coverage, segmentation defaults to the unsegmented word form itself. This resource is a necessary component for evaluating Tokenizer Morphological Plausibility, as introduced in Tokenizer Morphological Plausibility (https://arxiv.org/abs/2601.18536). The data generation process follows the implementation provided in the official… See the full description on the dataset page: https://huggingface.co/datasets/SHENJJ1017/morph_features.texttoken-classification10M<n<100M1 likes640 downloads8mo agoHugging Face27WecoAI /autodata-climbmix-features AutoData — ClimbMix Feature Bank Per-document annotations for the ClimbMix pre-training pool, used by the data-selection recipes in WecoAI/AutoData. Two banks are released: Directory Coverage Contents Size / (root) full pool — 553,155,584 docs lexical + perplexity + topic/format ~22.7 GiB reasoning_53shards/ first 4,485,120 docs Gemini reasoning/error annotations ~26 MiB License and provenance This dataset contains derived per-document… See the full description on the dataset page: https://huggingface.co/datasets/WecoAI/autodata-climbmix-features.100M<n<1B0 likes600 downloads19d agoHugging Face28FredZhang7 /malicious-website-features-2.4MImportant Notice: A subset of the URL dataset is from Kaggle, and the Kaggle datasets contained 10%-15% mislabelled data. See this dicussion I opened for some false positives. I have contacted Kaggle regarding their erroneous "Usability" score calculation for these unreliable datasets. The feature extraction methods shown here are not robust at all in 2023, and there're even silly mistakes in 3 functions: not_indexed_by_google, domain_registration_length, and age_of_domain. The features… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/malicious-website-features-2.4M.text-classification1M<n<10M6 likes598 downloads3y agoHugging Face29Peacockery /librispeech-phoneme-featurestabular100K<n<1M0 likes570 downloads7mo agoHugging Face30ozefe /spotify_audio_features Spotify Tracks & Audio Features Dataset Overview This dataset contains a comprehensive collection of Spotify tracks, combining rich audio feature analysis with track metadata. It is formatted as a high-performance Parquet dataset (ZStandard compressed), optimized for large-scale tabular analysis, machine learning, and recommender system research. Data Source The raw data for this dataset was originally gathered and hosted by Anna's Archive. Original Blog Post:… See the full description on the dataset page: https://huggingface.co/datasets/ozefe/spotify_audio_features.tabulartabular-regression100M<n<1B11 likes568 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.