Team Ai
Datasetpublic

Ian100/ProbeScout-features

ProbeScout features Frozen SigLIP features for Stanford Cars, HICO-DET and CelebA: float32 global embeddings [N, 768], float16 patch tokens [N, 196, 768], and records that map each row to an image ID. Total size: 81.10 GB. Use the ProbeScout setup guide and dataset-specific download commands. Download only the dataset you need. Place its dataset/ contents under the code repository's dataset/ directory, preserving paths and row order. These features skip extraction for training… See the full description on the dataset page: https://huggingface.co/datasets/Ian100/ProbeScout-features.

sourceHugging Faceupdated 19d agoView on Hugging Face
0likes326downloads
Dataset Card

ProbeScout features

Frozen SigLIP features for Stanford Cars, HICO-DET and CelebA: float32 global embeddings [N, 768], float16 patch tokens [N, 196, 768], and records that map each row to an image ID. Total size: 81.10 GB.

Use the ProbeScout setup guide and dataset-specific download commands. Download only the dataset you need. Place its dataset/ contents under the code repository's dataset/ directory, preserving paths and row order.

These features skip extraction for training and probe updates. Original images are still required for photograph previews. Browsing saved tasks and using the paper's Weight Tune / Staged feedback only require images and the task package.

CelebA patch tokens

The 60.99 GB CelebA patch NPY is stored as 29 byte-range parts. From the download directory containing restore_patch_features.py, run:

sh
python restore_patch_features.py --root .

Restoration verifies file hashes and needs an additional 60.99 GB of disk space. It keeps the parts. Cars and HICO patch tokens are ordinary NPY files.

The code pins payload revision 3ef237a82d0c738a159aa49bb86764732d331265. Checksums are in asset_manifest.json. Original dataset terms apply; photographs and historical human feedback are excluded.