keypa/vision-adapter-embeddings
Vision Adapter MoonViT Embeddings Precomputed visual embeddings used to train lightweight vision→LLM projectors without re-running a vision tower: each row is the frozen MoonViT-V2 output for one training image, stored as raw bfloat16 bytes. Shards: 103 Parquet files (data/emb_0000.parquet … data/emb_0102.parquet), 1360 rows each, ~139k rows total, ~0.9 TB. Schema per row: column type meaning key string embedding id, embeddings/<sha1[:20]>.pt; matches emb in… See the full description on the dataset page: https://huggingface.co/datasets/keypa/vision-adapter-embeddings.
Vision Adapter MoonViT Embeddings
Precomputed visual embeddings used to train lightweight vision→LLM projectors without re-running a vision tower: each row is the frozen MoonViT-V2 output for one training image, stored as raw bfloat16 bytes.
- Shards: 103 Parquet files (
data/emb_0000.parquet…data/emb_0102.parquet), 1360 rows each, ~139k rows total, ~0.9 TB. - Schema per row:
- Source images:
keypa/vision-adapter-images(MoonViT-V2 preprocessing: ≤300k pixels, 28-pixel multiples). Row key =sha1(path-after-images/)[:20]. - Row groups are small (≤128 rows, except
emb_0000/emb_0001which pack 1360 rows and are excluded from training plans as smoke shards). - Distribution is skewed: a 101–500 token majority plus a 4901+ long tail (measured p50 364, p99 5520, max 16653); training code buckets by
n_vis(0–100 / 101–500 / 501–1000 / 1001–2000 / 2001–4900 / 4901+).
n_vis is the ground truth for MoonViT geometry
MoonViT does not store its patch grid, and it cannot be recovered from a token count alone — n_vis=364 covers two genuinely different grids in this corpus. The grid is a deterministic function of the image dimensions under the preprocessing contract, so it was recovered from keypa/vision-adapter-images and written into train_manifest_grids.jsonl as a per-row grid_thw.
Every grid in that manifest reproduces the n_vis here — 117,600/117,600 verified against this dataset. If you recompute geometry, check it against this column rather than against a token count.
Intended use
Streaming access only — do not bulk-download. Reference implementation: vision_adapter/data/stream.py in keypa/Vision-Adapter (Range fetches + row-group cache + key index).
from datasets import load_dataset
ds = load_dataset("keypa/vision-adapter-embeddings", split="train", streaming=True)
row = next(iter(ds.select_columns(["key", "n_vis"])))
print(row["key"], row["n_vis"])