datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
frameflow-ltx-ae-vae-p256-full-2h-20261001
vae_p256_full_2h
Joint protein representation autoencoder ablation. See config.yaml, provenance.json and status.json for the exact configuration and progress. W&B: https://wandb.ai/gaorory-ucla-team/frameflow-ltx/runs/pr0o20u2
Exact FP32 model weights are retained every 5 optimizer steps inside immutable checkpoints/.tar shards. Each member is a torch checkpoint containing vae, config, step, and train_seconds. checkpoints.jsonl records member names and SHA-256 hashes. Extract a… See the full description on the dataset page: https://huggingface.co/datasets/raftbioworks/frameflow-ltx-ae-vae-p256-full-2h-20261001.frameflow-ltx-ae-vae-p128-full-2h-20261001
vae_p128_full_2h
Joint protein representation autoencoder ablation. See config.yaml, provenance.json and status.json for the exact configuration and progress. W&B: https://wandb.ai/gaorory-ucla-team/frameflow-ltx/runs/rhfznd0c
Exact FP32 model weights are retained every 5 optimizer steps inside immutable checkpoints/.tar shards. Each member is a torch checkpoint containing vae, config, step, and train_seconds. checkpoints.jsonl records member names and SHA-256 hashes. Extract a… See the full description on the dataset page: https://huggingface.co/datasets/raftbioworks/frameflow-ltx-ae-vae-p128-full-2h-20261001.frameflow-ltx-ae-vae-p64-full-2h-20261001
vae_p64_full_2h
Joint protein representation autoencoder ablation. See config.yaml, provenance.json and status.json for the exact configuration and progress. W&B: https://wandb.ai/gaorory-ucla-team/frameflow-ltx/runs/chify8j8
Exact FP32 model weights are retained every 5 optimizer steps inside immutable checkpoints/.tar shards. Each member is a torch checkpoint containing vae, config, step, and train_seconds. checkpoints.jsonl records member names and SHA-256 hashes. Extract a… See the full description on the dataset page: https://huggingface.co/datasets/raftbioworks/frameflow-ltx-ae-vae-p64-full-2h-20261001.frameflow-ltx-ae-vae-p512-full-2h-20261001
vae_p512_full_2h
Joint protein representation autoencoder ablation. See config.yaml, provenance.json and status.json for the exact configuration and progress. W&B: https://wandb.ai/gaorory-ucla-team/frameflow-ltx/runs/s9yuclhn
Exact FP32 model weights are retained every 5 optimizer steps inside immutable checkpoints/.tar shards. Each member is a torch checkpoint containing vae, config, step, and train_seconds. checkpoints.jsonl records member names and SHA-256 hashes. Extract a… See the full description on the dataset page: https://huggingface.co/datasets/raftbioworks/frameflow-ltx-ae-vae-p512-full-2h-20261001.frameflow-ltx-ae-vae-p512-data40k-2h-20261001
vae_p512_data40k_2h
Joint protein representation autoencoder ablation. See config.yaml, provenance.json and status.json for the exact configuration and progress. W&B: https://wandb.ai/gaorory-ucla-team/frameflow-ltx/runs/lfgj8krz
Exact FP32 model weights are retained every 5 optimizer steps inside immutable checkpoints/.tar shards. Each member is a torch checkpoint containing vae, config, step, and train_seconds. checkpoints.jsonl records member names and SHA-256 hashes. Extract… See the full description on the dataset page: https://huggingface.co/datasets/raftbioworks/frameflow-ltx-ae-vae-p512-data40k-2h-20261001.vae_cache_minecraft_480p_9sHDR_Photos_VAE_Training_DNGA collection of HDR images in the DNG format for use in training a HDR VAE.
frameflow-ltx-ae-vae-p512-full-12h-20261001
vae_p512_full_12h
Joint protein representation autoencoder ablation. See config.yaml, provenance.json and status.json for the exact configuration and progress. W&B: https://wandb.ai/gaorory-ucla-team/frameflow-ltx/runs/e7pjq3aj
Exact FP32 model weights are retained every 5 optimizer steps inside immutable checkpoints/.tar shards. Each member is a torch checkpoint containing vae, config, step, and train_seconds. checkpoints.jsonl records member names and SHA-256 hashes. Extract a… See the full description on the dataset page: https://huggingface.co/datasets/raftbioworks/frameflow-ltx-ae-vae-p512-full-12h-20261001.frameflow-ltx-ae-vae-p512-full-6h-20261001
vae_p512_full_6h
Joint protein representation autoencoder ablation. See config.yaml, provenance.json and status.json for the exact configuration and progress. W&B: https://wandb.ai/gaorory-ucla-team/frameflow-ltx/runs/mmcwntus
Exact FP32 model weights are retained every 5 optimizer steps inside immutable checkpoints/.tar shards. Each member is a torch checkpoint containing vae, config, step, and train_seconds. checkpoints.jsonl records member names and SHA-256 hashes. Extract a… See the full description on the dataset page: https://huggingface.co/datasets/raftbioworks/frameflow-ltx-ae-vae-p512-full-6h-20261001.pdm3-ht-20260528-flux2-vae-latents-public
PDM-3-HT FLUX.2 VAE latents for ImageNet-256 train
Public research artifact for PDM-3-HT VAE-backend experiments. This repository contains latent cache shards only. It intentionally does not contain raw ImageNet images, ADM-cropped uint8 images, PAE latents, PAE checkpoints, or training checkpoints.
Source and preprocessing
Source dataset: ImageNet-1k train via ILSVRC/imagenet-1k; access requires accepting the upstream ImageNet terms.
Image preprocessing before VAE… See the full description on the dataset page: https://huggingface.co/datasets/LAXMAYDAY/pdm3-ht-20260528-flux2-vae-latents-public.font-square-v2-pairs-vaeimagenet_vae_mds_fp32vaeqythr-0
Sharp Direct-Sum Savings in Sparse Positive-Semidefinite Factorizations
VAEQYTHR-0 · version 3.0.0 · 4 October 2026
This research-data repository contains a mathematical manuscript, exact matrices,
executable certificates, and proof-audit records. Its subject is the number of
sparse rank-one terms needed to represent a positive-semidefinite matrix.
It contains no trained model, training corpus, or machine-learning benchmark.
Manuscript: PDF ·
LaTeX source ·
Theorem ledger… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/vaeqythr-0.frameflow-ltx-joint-vae-20261001
FrameFlow × LTX: joint_vae
Research experiment adapting LTX-Video, especially Figure 4,
to representations from vanilla FrameFlow.
This is a protein adaptation, not a reproduction of the video model.
Training completed 419,933 optimizer updates: 2.000 VAE training hours
and 12.000 latent-flow training hours.
The codec is frozen and its latents calibrated on training examples before flow training.
W&B run.
Contents
All interval/final checkpoints with SHA-256 files;… See the full description on the dataset page: https://huggingface.co/datasets/raftbioworks/frameflow-ltx-joint-vae-20261001.imagenet-256-flux2-vae-latents
ImageNet-256 FLUX.2 VAE Latents
Pre-computed deterministic, model-facing encodings from the
FLUX.2 VAE (black-forest-labs/FLUX.2-dev)
for the full ImageNet-1K training set at 256x256 resolution, stored as Parquet
shards. Each example includes latents for both the original and horizontally
flipped image, enabling flip augmentation without re-encoding at training time.
Dataset Description
Each example contains:
Column
Shape
Stored type
Description… See the full description on the dataset page: https://huggingface.co/datasets/yuanchenyang/imagenet-256-flux2-vae-latents.Audio-VAE-Phonk-Dataset
🚗 Phonk Audio Dataset for Generative ML
This dataset contains hundreds of hours of high-quality Phonk music (Drift Phonk, Hard Phonk, etc.), specifically scraped, pre-processed, and formatted for training deep learning audio models.
It is perfectly suited for training Audio VAEs, EnCodec, TiTok, or Audio Diffusion / Transformer Prior models from scratch.
📊 Dataset Specifications
The audio data has been heavily pre-processed to maximize training efficiency on TPUs/GPUs:… See the full description on the dataset page: https://huggingface.co/datasets/Prhokbvf556/Audio-VAE-Phonk-Dataset.frameflow-ltx-pair-vae-20261001
FrameFlow × LTX: pair_vae
Research experiment adapting LTX-Video, especially Figure 4,
to representations from vanilla FrameFlow.
This is a protein adaptation, not a reproduction of the video model.
Training completed 451,619 optimizer updates: 2.000 VAE training hours
and 12.000 latent-flow training hours.
The codec is frozen and its latents calibrated on training examples before flow training.
W&B run.
Contents
All interval/final checkpoints with SHA-256 files;… See the full description on the dataset page: https://huggingface.co/datasets/raftbioworks/frameflow-ltx-pair-vae-20261001.cache_vae_latents_train_20hq1lq-part-00imagenet-latents-qwen-image-vaevaeframeflow-ltx-ae-vae-p512-data10k-2h-20261001
vae_p512_data10k_2h
Joint protein representation autoencoder ablation. See config.yaml, provenance.json and status.json for the exact configuration and progress. W&B: https://wandb.ai/gaorory-ucla-team/frameflow-ltx/runs/258co4h7
Exact FP32 model weights are retained every 5 optimizer steps inside immutable checkpoints/.tar shards. Each member is a torch checkpoint containing vae, config, step, and train_seconds. checkpoints.jsonl records member names and SHA-256 hashes. Extract… See the full description on the dataset page: https://huggingface.co/datasets/raftbioworks/frameflow-ltx-ae-vae-p512-data10k-2h-20261001.laion-pop-vae-t5mix_vae_tensorVAEscapstone_sakuga_vae_latentsqdump_s2_qrep_vae128_step021000imagenet-sdxl-vae-uint8curated-danbooru-2026-512px-flux2-vaeimagenet-latents-qwenimage-vae-e2e-lr2e-5-400kframeflow-ltx-residue-vae-20261001
FrameFlow × LTX: residue_vae
Research experiment adapting LTX-Video, especially Figure 4,
to representations from vanilla FrameFlow.
This is a protein adaptation, not a reproduction of the video model.
Training completed 433,004 optimizer updates: 2.000 VAE training hours
and 12.000 latent-flow training hours.
The codec is frozen and its latents calibrated on training examples before flow training.
W&B run.
Contents
All interval/final checkpoints with SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/raftbioworks/frameflow-ltx-residue-vae-20261001.
