Team Ai
Datasetpublic

keypa/vision-adapter-embeddings

Vision Adapter MoonViT Embeddings Precomputed visual embeddings used to train lightweight vision→LLM projectors without re-running a vision tower: each row is the frozen MoonViT-V2 output for one training image, stored as raw bfloat16 bytes. Shards: 103 Parquet files (data/emb_0000.parquet … data/emb_0102.parquet), 1360 rows each, ~139k rows total, ~0.9 TB. Schema per row: column type meaning key string embedding id, embeddings/<sha1[:20]>.pt; matches emb in… See the full description on the dataset page: https://huggingface.co/datasets/keypa/vision-adapter-embeddings.

sourceHugging Faceunknownupdated 5d agoView on Hugging Face
0likes2kdownloads
Dataset Card

Vision Adapter MoonViT Embeddings

Precomputed visual embeddings used to train lightweight vision→LLM projectors without re-running a vision tower: each row is the frozen MoonViT-V2 output for one training image, stored as raw bfloat16 bytes.

  • —Shards: 103 Parquet files (data/emb_0000.parquet … data/emb_0102.parquet), 1360 rows each, ~139k rows total, ~0.9 TB.
  • —Schema per row:
columntypemeaning
keystringembedding id, embeddings/<sha1[:20]>.pt; matches emb in keypa/vision-adapter-manifests/train_manifest_grids.jsonl
n_visint64number of visual tokens in this row (16 … ~16653)
vis_bytesbinaryn_vis × 4096 bfloat16 little-endian (torch.from_numpy(u8).view(bf16).reshape(-1, 4096))
  • —Source images: keypa/vision-adapter-images (MoonViT-V2 preprocessing: ≤300k pixels, 28-pixel multiples). Row key = sha1(path-after-images/)[:20].
  • —Row groups are small (≤128 rows, except emb_0000/emb_0001 which pack 1360 rows and are excluded from training plans as smoke shards).
  • —Distribution is skewed: a 101–500 token majority plus a 4901+ long tail (measured p50 364, p99 5520, max 16653); training code buckets by n_vis (0–100 / 101–500 / 501–1000 / 1001–2000 / 2001–4900 / 4901+).

n_vis is the ground truth for MoonViT geometry

MoonViT does not store its patch grid, and it cannot be recovered from a token count alone — n_vis=364 covers two genuinely different grids in this corpus. The grid is a deterministic function of the image dimensions under the preprocessing contract, so it was recovered from keypa/vision-adapter-images and written into train_manifest_grids.jsonl as a per-row grid_thw.

Every grid in that manifest reproduces the n_vis here — 117,600/117,600 verified against this dataset. If you recompute geometry, check it against this column rather than against a token count.

Intended use

Streaming access only — do not bulk-download. Reference implementation: vision_adapter/data/stream.py in keypa/Vision-Adapter (Range fetches + row-group cache + key index).

python
from datasets import load_dataset
ds = load_dataset("keypa/vision-adapter-embeddings", split="train", streaming=True)
row = next(iter(ds.select_columns(["key", "n_vis"])))
print(row["key"], row["n_vis"])