Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SwayStar123 /preprocessed_commoncatalog-cc-byI also seperately provide just the prompts in prompts.json keys are the image_id, and the values are the captions generated Captions generated by moondream: vikhyatk/moondream2 Latents generated by SDXL VAE: madebyollin/sdxl-vae-fp16-fix Embeddings generated by SigLIP: hf-hub:timm/ViT-SO400M-14-SigLIP-384 Original dataset: common-canvas/commoncatalog-cc-by Latents f32 and embeddings are f16 bytes Compute cost: 16x3090 for 3 day. Approximately. text10M<n<100M19 likes493k downloads2y agoHugging Face02SwayStar123 /preprocessed_commoncatalog-cc-by_DCAEThe images are resized and then encoded with the DC-AE f32 autoencoder. The resizing is done with a bucketmanager with base resolution 512x512, minimum side length 256, maximum side length 1024, all sides are divisible by 32 ofcourse as they needed to be encoded by the DCAEf32 encoder. The captions are generated with moondream2, encoded with siglip and bert. (Bert embeddings variance is very high, so use a norm layer). The text embeddings are padded to 64 tokens, but i have provided the… See the full description on the dataset page: https://huggingface.co/datasets/SwayStar123/preprocessed_commoncatalog-cc-by_DCAE.text-to-image10M<n<100M1 likes36k downloads2y agoHugging Face03SwayStar123 /preprocessed_DCAE-f64_1024_commoncatalog-cc-bytext10M<n<100M0 likes6.5k downloads2y agoHugging Face04Jang-Hyun /SCBench-preprocessedThis is the preprocessed version of Microsoft SCBench, used by KVzip: Each data example has a format of {context: str, question: List[str], answers: List[str]} Each dataset contains only examples whose context token length (measured with the LLaMA3 tokenizer) is less than 125K, fitting within the context limit of LLaMA3 models. We also provide shortened versions of SCBench, excluding tasks {choice_eng, qa_eng, and vt}, which are difficult to shorten. The "tiny" tag (e.g., scbench_kv_tiny)… See the full description on the dataset page: https://huggingface.co/datasets/Jang-Hyun/SCBench-preprocessed.text1K<n<10K2 likes5.3k downloads9mo agoHugging Face05stanford-star /relbench-preprocessed RelBench, preprocessed stanford-star/relbench-v1 in the tensor format read by the Relational Transformer: 7 databases, 21 forecast and 13 autocomplete tasks. The evaluation and validation data for RT-J and RT-PluRel. One directory per database: <db>/ meta.json table_info.json column_index.json nodes.rkyv offsets.rkyv p2f_adj.rkyv text.json text_emb_all-MiniLM-L12-v2.bin legacy/ holds the same databases preprocessed with the boolean typing expected by the… See the full description on the dataset page: https://huggingface.co/datasets/stanford-star/relbench-preprocessed.tabulartabular-classification10M<n<100M0 likes4.1k downloads6d agoHugging Face06Yinpei /robomme_preprocessed_data RoboMME Training Data (Pickle Format) Arxiv Paper | HF Paper | Website | Benchmark Code | Policy Learning Code This repo contains preprocessed pickle files for RoboMME training data and npy files for cached image tokens. We use this dataset in our MME-VLA experiments. . ├── data # zipped pickle files ├── features # zipped precompute siglip embeddings ├── meta # statistics for robomme ├── memer # VLM subgoal training data for MemER (only used for symbolic… See the full description on the dataset page: https://huggingface.co/datasets/Yinpei/robomme_preprocessed_data.image100K<n<1M0 likes3.2k downloads7mo agoHugging Face07stanford-star /the-join-preprocessed The Join, preprocessed stanford-star/the-join in the tensor format read by the Relational Transformer: the 523 databases whose raw tables are under 5 GiB, 13,243 tasks. Phase 2 of RT-J pretraining. One directory per database: <db>/ meta.json table_info.json column_index.json nodes.rkyv offsets.rkyv p2f_adj.rkyv text.json text_emb_all-MiniLM-L12-v2.bin Download, subset and revision pins: examples/README.md. Preprocessing: examples/preprocess/.… See the full description on the dataset page: https://huggingface.co/datasets/stanford-star/the-join-preprocessed.tabulartabular-classification100B<n<1T0 likes3.2k downloads6d agoHugging Face08KevinConnorLee /complet4r_preprocessed_sailvos3d0 likes2.5k downloads9mo agoHugging Face09cyn-xyz /ideal_scenes_preprocessed0 likes2.3k downloads6mo agoHugging Face10SwayStar123 /preprocessed_DCAE-f64_1024_pd12m-fulltext1M<n<10M0 likes2.1k downloads2y agoHugging Face11collabora /hi-stt-preprocessed-webdatasettext100K<n<1M1 likes2.1k downloads1y agoHugging Face12HayrettinIscan /MeshAI-Preprocessed-4Kimage0 likes1.9k downloads3mo agoHugging Face13stanford-star /plurel-preprocessed PluRel, preprocessed stanford-star/plurel in the tensor format read by the Relational Transformer: 2,000 synthetic databases, plurel-3000 to plurel-4999. Pretraining corpus of RT-PluRel and phase 1 of RT-J, which train on the 86,211 tasks over 1,900 databases selected by rt.data.plurel_train_db_task_list. One directory per database: <db>/ meta.json table_info.json column_index.json nodes.rkyv offsets.rkyv p2f_adj.rkyv text.json text_emb_all-MiniLM-L12-v2.bin… See the full description on the dataset page: https://huggingface.co/datasets/stanford-star/plurel-preprocessed.tabulartabular-classification10B<n<100B0 likes1.1k downloads6d agoHugging Face14tankalapavankalyan /ds007808-sub01-listening-pangolin-preprocessed0 likes1.1k downloads4mo agoHugging Face15INo0121 /low_quality_call_voice_preprocessed Dataset Card for "low_quality_call_voice_preprocessed" More Information needed 10K<n<100K1 likes1k downloads3y agoHugging Face16KevinConnorLee /complet4r_preprocessed_dynamicreplica0 likes970 downloads4mo agoHugging Face17ENSEONG /preprocessed-full-math-private-n256-Llama-3.2-3B-Instruct-bontabular100K<n<1M0 likes870 downloads5mo agoHugging Face18verl-team /lighteval-MATH-preprocessedtext10K<n<100K1 likes867 downloads2y agoHugging Face19ddPn08 /maestro-v3.0.0-preprocessed0 likes807 downloads2y agoHugging Face20jpbello /common_language_preprocessed Dataset Card for "common_language_preprocessed" More Information needed text10K<n<100K0 likes798 downloads3y agoHugging Face21SwayStar123 /preprocessed_DCAE-f64_commoncatalog-cc-by-satext1M<n<10M0 likes720 downloads2y agoHugging Face22maxseats /aihub-464-preprocessed-680GB-set-52audio10K<n<100K0 likes711 downloads2y agoHugging Face23SwayStar123 /preprocessed_DCAE-f64_commoncatalog-cc-bytext10M<n<100M0 likes609 downloads2y agoHugging Face24Self-GRIT /wikitext-2-raw-v1-preprocessedtext10K<n<100K2 likes603 downloads2y agoHugging Face25Rubin-Wei /enwiki-dec2021-preprocessed-mistral Dataset Description This dataset is a preprocessed version of the English Wikipedia snapshot from December 2021, processed using the preprocess_dataset.py script provided in the repository below. Paper: MLP Memory: A Retriever-Pretrained Memory for Large Language Models GitHub: https://github.com/Rubin-Wei/MLPMemory Dataset Source: English Wikipedia (December 2021) Tokenizer: Mistral-7B-v0.3 Two key preprocessing parameters used are: block_size: 2048 stride: 1024… See the full description on the dataset page: https://huggingface.co/datasets/Rubin-Wei/enwiki-dec2021-preprocessed-mistral.1M<n<10M0 likes598 downloads1y agoHugging Face26chaenayo /nabirds_custom_split_preprocessedimage10K<n<100K0 likes593 downloads1y agoHugging Face27tankalapavankalyan /ds007808-sub01-covert-pangolin-preprocessed0 likes582 downloads4mo agoHugging Face28JasonGao726 /ISPRS_Potsdam_dataset_preprocessed0 likes579 downloads26d agoHugging Face29hourouu /LUNA16_preprocessed100K<n<1M0 likes539 downloads5mo agoHugging Face30makaveli10 /whisper-hi-preprocessed Dataset Card for "whisper-hi-preprocessed" More Information needed 1K<n<10K0 likes521 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.