ByteDance/Bernini-R
29991
1---2license: apache-2.03pipeline_tag: image-text-to-video4---5<div align="center">6 7<img src="assets/bernini-icon.png" width="560" alt="Bernini"/>8 9<h4 align="center">Latent Semantic Planning for Video Diffusion</h4>10 11**Chenchen Liu<sup>\*</sup>, Junyi Chen<sup>\*</sup>, Lei Li<sup>\*</sup>, Lu Chi<sup>\*,§</sup>, Mingzhen Sun<sup>\*</sup>, Zhuoying Li<sup>\*</sup>, Yi Fu, Ruoyu Guo, Yiheng Wu, Ge Bai, Zehuan Yuan<sup>✉</sup>**12 13<sup>\*</sup> Equal contribution <sup>✉</sup> Corresponding author <sup>§</sup> Project lead14 15[](https://arxiv.org/abs/2605.22344)16[](https://bernini-ai.github.io/)17[](https://huggingface.co/ByteDance/Bernini)18 19</div>20 21## 🎉 News22 23- **[2026-06-01]** We open-sourced the inference code and model weights of the Bernini Renderer (**Bernini-R**).24- **[2026-05-22]** We released our paper [Bernini: Latent Semantic Planning for Video Diffusion](https://arxiv.org/abs/2605.22344).25 26## ✨ Highlights27 28Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.29 30On video editing, Bernini reaches the first tier among leading closed-source31commercial models. The leaderboard below comes from our self-built arena32platform, where human annotators blindly vote on paired edits and the votes are33aggregated into a Bradley-Terry score and a pairwise win-rate matrix.34 35<img src="assets/arena.png" width="900" alt="Video editing arena: Bradley-Terry leaderboard and pairwise win-rate matrix"/>36 37## 📦 Installation38 39### Requirements40 41- **Python** 3.11.2.42- **CUDA GPU** — a Hopper GPU (H100/H800/H200) is recommended so FlashAttention-343 can be used; other CUDA GPUs fall back to FlashAttention-2 or PyTorch SDPA.44- **CUDA toolkit** 12.4 (matches the pinned `torch==2.5.1+cu124`; 12.3+ is the45 minimum if you build FlashAttention-3).46- Pinned in `requirements.txt`: `torch==2.5.1+cu124`, `diffusers==0.35.2`,47 `accelerate==0.34.2`, `transformers==4.57.3`.48 49Reference environment (Bernini-R is developed and tested on this setup):50 51| Component | Version |52|-----------|--------------|53| GPU | NVIDIA H100 |54| CUDA | 12.4 |55| Python | 3.11.2 |56| PyTorch | 2.5.1+cu124 |57 58### Install59 60```bash61git clone https://github.com/bytedance/Bernini.git bernini && cd bernini62pip install -r requirements.txt63```64 65Optional extras:66 67- **Multi-GPU sequence parallel** needs [Open-VeOmni](https://github.com/ByteDance-Seed/VeOmni)68 (Apache-2.0, Python 3.11). Use `--no-deps` so VeOmni does not pull in a69 different torch build and override the pinned `torch==2.5.1+cu124`:70 `pip install --no-deps git+https://github.com/ByteDance-Seed/VeOmni.git@v0.1.10`.71 Single-GPU inference does not need it.72- **Faster attention** (auto-detected if installed; otherwise PyTorch SDPA is used):73 - FlashAttention-2 — general CUDA GPUs (incl. A100/A800): `pip install flash-attn==2.8.3`.74 - FlashAttention-3 — Hopper only (H100/H800/H200, CUDA ≥ 12.3, PyTorch ≥ 2.4).75 `flash_attn_interface` is not on PyPI; build it from the76 [flash-attention](https://github.com/Dao-AILab/flash-attention) repo's77 `hopper/` directory at tag `v2.8.3`:78 ```bash79 git clone https://github.com/Dao-AILab/flash-attention.git80 cd flash-attention && git checkout v2.8.381 cd hopper && MAX_JOBS=$(nproc) python3 setup.py install --user82 ```83 84### Weights85 86Bernini-R provides two ways to obtain the renderer weights. The **diffusers87format is recommended** — it is a self-contained diffusers-format directory whose88`transformer` / `transformer_2` already hold the Bernini-R weights, so you point89`--config` at it and the weights load directly, with **no** `--high_noise_ckpt` /90`--low_noise_ckpt` needed.91 92#### Option A — diffusers format (recommended)93 94A single ready-to-use diffusers-format model from95[`ByteDance/Bernini-R-Diffusers`](https://huggingface.co/ByteDance/Bernini-R-Diffusers).96It bundles the Wan2.2 base components (VAE, UMT5 text encoder, tokenizer) together97with the Bernini-R transformer weights, so nothing else is downloaded at runtime.98 99```bash100pip install -U "huggingface_hub"101hf download ByteDance/Bernini-R-Diffusers --local-dir Bernini-R-Diffusers102```103 104Then pass it via `--config` and omit the checkpoint flags, e.g.:105 106```bash107python infer_single_gpu.py --config Bernini-R-Diffusers \108 --case assets/testcases/t2i/t2i.json --num_frames 1109```110 111#### Option B — separate checkpoints112 113The original layout, where Bernini-R uses two sets of weights loaded separately:114 1151. **Wan2.2 base** — [`Wan-AI/Wan2.2-T2V-A14B-Diffusers`](https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B-Diffusers) on Hugging Face. Supplies the116 VAE, UMT5 text encoder, tokenizer, and the transformer architecture/base weights.117 It is downloaded automatically on first run (configured by `wan22_base` in118 `configs/bernini_renderer_wan22/config.json`).1192. **Bernini-R checkpoint** — the trained high-noise / low-noise transformer weights120 (safetensors) from [ByteDance/Bernini-R](https://huggingface.co/ByteDance/Bernini-R), passed with121 `--high_noise_ckpt` / `--low_noise_ckpt`. Both a local directory and a Hugging122 Face repo id are accepted.123 124Download models using huggingface-cli:125 126```bash127pip install -U "huggingface_hub"128hf download Wan-AI/Wan2.2-T2V-A14B-Diffusers --local-dir Wan2.2-T2V-A14B-Diffusers129hf download ByteDance/Bernini-R --local-dir Bernini-R130```131 132## 🚀 Usage133 134A run is described by a **case file** — a small JSON under135[`assets/testcases/`](assets/testcases/) that bundles one task's routing and136inputs (`task_type`, `guidance_mode`, `prompt`, source media, `output`). This137keeps long prompts out of the command line. Each task has a directory under138`assets/testcases/` holding one or more case files; see139[`assets/testcases/`](assets/testcases/) for the format and the bundled140`t2i` / `i2i` / `t2v` / `v2v` / `rv2v` /`r2v` examples.141 142### Prompt enhancer (highly recommended)143 144`--use_pe` enhances the prompt through an OpenAI-compatible endpoint and is145recommended for best generation quality. The `openai` SDK is installed by146`requirements.txt`; configure the endpoint with environment variables:147 148```bash149export BERNINI_PE_API_KEY=... # or OPENAI_API_KEY150export BERNINI_PE_BASE_URL=... # or OPENAI_BASE_URL151export BERNINI_PE_MODEL=... # vision-capable chat model152```153 154### Examples by task type155 156Unless an example specifies otherwise, inference outputs **480p / 16fps** (the157defaults — `--max_image_size 848`, `--fps 16`).158 159Each example runs a bundled case in160[`assets/testcases/`](assets/testcases/) — replace `<hi>` / `<lo>` with your161high-/low-noise checkpoint paths. The image tasks (`t2i`, `i2i`) are shown on a162single GPU; the video tasks on 8 GPUs via `torchrun`, where `--ulysses N` gives163N-way Ulysses sequence parallel per sample and the remaining `world_size / N`164ranks run data parallel over the task list. The two scripts take the same165inputs, so any example can be run either way.166 167Inputs can also be passed directly as flags instead of `--case` (`--prompt`,168`--task_type`, `--guidance_mode`, `--video`, `--image`, `--images`,169`--output`); generation parameters (`--seed`, `--num_frames`, ...) are always170command-line flags.171 172**Text-to-image** (`t2i`) — single GPU; generates one frame, so pass `--num_frames 1`173 174```bash175python infer_single_gpu.py --high_noise_ckpt <hi> --low_noise_ckpt <lo> \176 --case assets/testcases/t2i/t2i.json --num_frames 1177```178 179**Image editing** (`i2i`) — single GPU; generates one frame, so pass `--num_frames 1`180 181```bash182python infer_single_gpu.py --high_noise_ckpt <hi> --low_noise_ckpt <lo> \183 --case assets/testcases/i2i/i2i.json --num_frames 1184```185 186**Text-to-video** (`t2v`)187 188```bash189torchrun --nproc-per-node 8 infer_multi_gpu.py \190 --high_noise_ckpt <hi> --low_noise_ckpt <lo> --ulysses 8 \191 --case assets/testcases/t2v/t2v.json192```193 194**Video editing** (`v2v` / `mv2v`) — two cases are provided.195 196For edits where the main subject keeps its ordinary motion (case 1 adds a197snowman to the scene), the `v2v` task type is enough:198 199```bash200torchrun --nproc-per-node 8 infer_multi_gpu.py \201 --high_noise_ckpt <hi> --low_noise_ckpt <lo> --ulysses 8 \202 --case assets/testcases/v2v/v2v_case1.json203```204 205For edits that need to change the subject's motion (case 2 makes the person206crouch down), the `mv2v` task type gives better results:207 208```bash209torchrun --nproc-per-node 8 infer_multi_gpu.py \210 --high_noise_ckpt <hi> --low_noise_ckpt <lo> --ulysses 8 \211 --case assets/testcases/v2v/v2v_case2.json212```213 214**Reference + video editing** (`rv2v`) — two cases are provided.215 216Case 1 is reference-image-guided video editing — replacing a garment in the217source video with one from a reference image:218 219```bash220torchrun --nproc-per-node 8 infer_multi_gpu.py \221 --high_noise_ckpt <hi> --low_noise_ckpt <lo> --ulysses 8 \222 --case assets/testcases/rv2v/rv2v_case1.json223```224 225Case 2 is a video-insertion example — inserting content into the source video.226It is run at 720p / 24fps to show the insertion result more clearly:227 228```bash229torchrun --nproc-per-node 8 infer_multi_gpu.py \230 --high_noise_ckpt <hi> --low_noise_ckpt <lo> --ulysses 8 \231 --case assets/testcases/rv2v/rv2v_case2.json \232 --num_frames 121 --fps 24 --max_image_size 1280233```234 235**Reference-to-video** (`r2v`) — drives a video from one or more reference images236 237```bash238torchrun --nproc-per-node 8 infer_multi_gpu.py \239 --high_noise_ckpt <hi> --low_noise_ckpt <lo> --ulysses 8 \240 --case assets/testcases/r2v/r2v.json241```242 243See `python infer_single_gpu.py --help` for the full argument list.244 245### Gradio demo246 247`gradio_demo.py` exposes the same pipeline through a Gradio UI: the task-type248dropdown auto-fills `guidance_mode` (still user-editable), uploaded media is249routed to the matching slot, and the result is rendered inline.250 251```bash252# Single GPU253python gradio_demo.py --high_noise_ckpt <hi> --low_noise_ckpt <lo> --port 7860254 255# 8 GPUs, 8-way Ulysses sequence parallel256torchrun --nproc-per-node 8 gradio_demo.py --ulysses 8 \257 --high_noise_ckpt <hi> --low_noise_ckpt <lo> --port 7860 --share258```259 260Add `--use_pe` (and `export OPENAI_API_KEY=...` / `BERNINI_PE_API_KEY=...`) to261enable GPT prompt enhancement; the in-UI checkbox is a per-request switch on262top of this flag.263 264## 📑 Citation265 266If you use Bernini in your research, please cite:267 268```bibtex269@article{bernini,270 title = {Bernini: Latent Semantic Planning for Video Diffusion},271 author = {Chenchen Liu and Junyi Chen and Lei Li and Lu Chi and Mingzhen Sun and Zhuoying Li and Yi Fu and Ruoyu Guo and Yiheng Wu and Ge Bai and Zehuan Yuan},272 journal = {arXiv preprint arXiv:2605.22344},273 year = {2026}274}275```276 277## 🙏 Acknowledgements278 279Bernini builds on several outstanding open-source projects:280 281- [Wan2.2-T2V-A14B](https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B)282- [Qwen2.5-VL-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-VL-7B-Instruct)283- [VeOmni](https://github.com/ByteDance-Seed/VeOmni)284 285We thank the authors and communities of these projects for their contributions.286 287## 📄 License288 289Apache License 2.0. See [LICENSE](LICENSE).