datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
preprocessed_commoncatalog-cc-byI also seperately provide just the prompts in prompts.json
keys are the image_id, and the values are the captions generated
Captions generated by moondream: vikhyatk/moondream2
Latents generated by SDXL VAE: madebyollin/sdxl-vae-fp16-fix
Embeddings generated by SigLIP: hf-hub:timm/ViT-SO400M-14-SigLIP-384
Original dataset: common-canvas/commoncatalog-cc-by
Latents f32 and embeddings are f16 bytes
Compute cost: 16x3090 for 3 day. Approximately.
preprocessed_commoncatalog-cc-by_DCAEThe images are resized and then encoded with the DC-AE f32 autoencoder. The resizing is done with a bucketmanager with base resolution 512x512, minimum side length 256, maximum side length 1024, all sides are divisible by 32 ofcourse as they needed to be encoded by the DCAEf32 encoder.
The captions are generated with moondream2, encoded with siglip and bert. (Bert embeddings variance is very high, so use a norm layer). The text embeddings are padded to 64 tokens, but i have provided the… See the full description on the dataset page: https://huggingface.co/datasets/SwayStar123/preprocessed_commoncatalog-cc-by_DCAE.preprocessed_DCAE-f64_1024_commoncatalog-cc-bySCBench-preprocessedThis is the preprocessed version of Microsoft SCBench, used by KVzip:
Each data example has a format of {context: str, question: List[str], answers: List[str]}
Each dataset contains only examples whose context token length (measured with the LLaMA3 tokenizer) is less than 125K, fitting within the context limit of LLaMA3 models.
We also provide shortened versions of SCBench, excluding tasks {choice_eng, qa_eng, and vt}, which are difficult to shorten.
The "tiny" tag (e.g., scbench_kv_tiny)… See the full description on the dataset page: https://huggingface.co/datasets/Jang-Hyun/SCBench-preprocessed.relbench-preprocessed
RelBench, preprocessed
stanford-star/relbench-v1 in the tensor format read by the Relational Transformer: 7 databases, 21 forecast and 13 autocomplete tasks. The evaluation and validation data for RT-J and RT-PluRel. One directory per database:
<db>/ meta.json table_info.json column_index.json
nodes.rkyv offsets.rkyv p2f_adj.rkyv
text.json text_emb_all-MiniLM-L12-v2.bin
legacy/ holds the same databases preprocessed with the boolean typing expected by the… See the full description on the dataset page: https://huggingface.co/datasets/stanford-star/relbench-preprocessed.robomme_preprocessed_data
RoboMME Training Data (Pickle Format)
Arxiv Paper | HF Paper | Website | Benchmark Code | Policy Learning Code
This repo contains preprocessed pickle files for RoboMME training data and npy files for cached image tokens. We use this dataset in our MME-VLA experiments.
.
├── data # zipped pickle files
├── features # zipped precompute siglip embeddings
├── meta # statistics for robomme
├── memer # VLM subgoal training data for MemER (only used for symbolic… See the full description on the dataset page: https://huggingface.co/datasets/Yinpei/robomme_preprocessed_data.the-join-preprocessed
The Join, preprocessed
stanford-star/the-join in the tensor format read by the Relational Transformer: the 523 databases whose raw tables are under 5 GiB, 13,243 tasks. Phase 2 of RT-J pretraining. One directory per database:
<db>/ meta.json table_info.json column_index.json
nodes.rkyv offsets.rkyv p2f_adj.rkyv
text.json text_emb_all-MiniLM-L12-v2.bin
Download, subset and revision pins: examples/README.md. Preprocessing: examples/preprocess/.… See the full description on the dataset page: https://huggingface.co/datasets/stanford-star/the-join-preprocessed.complet4r_preprocessed_sailvos3dideal_scenes_preprocessedpreprocessed_DCAE-f64_1024_pd12m-fullhi-stt-preprocessed-webdatasetMeshAI-Preprocessed-4Kplurel-preprocessed
PluRel, preprocessed
stanford-star/plurel in the tensor format read by the Relational Transformer: 2,000 synthetic databases, plurel-3000 to plurel-4999. Pretraining corpus of RT-PluRel and phase 1 of RT-J, which train on the 86,211 tasks over 1,900 databases selected by rt.data.plurel_train_db_task_list. One directory per database:
<db>/ meta.json table_info.json column_index.json
nodes.rkyv offsets.rkyv p2f_adj.rkyv
text.json text_emb_all-MiniLM-L12-v2.bin… See the full description on the dataset page: https://huggingface.co/datasets/stanford-star/plurel-preprocessed.ds007808-sub01-listening-pangolin-preprocessedlow_quality_call_voice_preprocessed
Dataset Card for "low_quality_call_voice_preprocessed"
More Information needed
complet4r_preprocessed_dynamicreplicapreprocessed-full-math-private-n256-Llama-3.2-3B-Instruct-bonlighteval-MATH-preprocessedmaestro-v3.0.0-preprocessedcommon_language_preprocessed
Dataset Card for "common_language_preprocessed"
More Information needed
preprocessed_DCAE-f64_commoncatalog-cc-by-saaihub-464-preprocessed-680GB-set-52preprocessed_DCAE-f64_commoncatalog-cc-bywikitext-2-raw-v1-preprocessedenwiki-dec2021-preprocessed-mistral
Dataset Description
This dataset is a preprocessed version of the English Wikipedia snapshot from December 2021, processed using the preprocess_dataset.py script provided in the repository below.
Paper: MLP Memory: A Retriever-Pretrained Memory for Large Language Models
GitHub: https://github.com/Rubin-Wei/MLPMemory
Dataset Source: English Wikipedia (December 2021)
Tokenizer: Mistral-7B-v0.3
Two key preprocessing parameters used are:
block_size: 2048
stride: 1024… See the full description on the dataset page: https://huggingface.co/datasets/Rubin-Wei/enwiki-dec2021-preprocessed-mistral.nabirds_custom_split_preprocessedds007808-sub01-covert-pangolin-preprocessedISPRS_Potsdam_dataset_preprocessedLUNA16_preprocessedwhisper-hi-preprocessed
Dataset Card for "whisper-hi-preprocessed"
More Information needed
