Team Ai
Modelpublic

ByteDance/Bernini-R-1.3B-Diffusers

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
28likes
README.md297 linesDownload Raw Back to root
1---2license: apache-2.03pipeline_tag: image-text-to-video4---5<div align="center">6 7<img src="assets/bernini-icon.png" width="560" alt="Bernini"/>8 9<h4 align="center">Latent Semantic Planning for Video Diffusion</h4>10 11**Chenchen Liu<sup>\*</sup>, Junyi Chen<sup>\*</sup>, Lei Li<sup>\*</sup>, Lu Chi<sup>\*,§</sup>, Mingzhen Sun<sup>\*</sup>, Zhuoying Li<sup>\*</sup>, Yi Fu, Ruoyu Guo, Yiheng Wu, Ge Bai, Zehuan Yuan<sup>✉</sup>**12 13<sup>\*</sup> Equal contribution&nbsp;&nbsp;<sup>✉</sup> Corresponding author&nbsp;&nbsp;<sup>§</sup> Project lead14 15[![arXiv](https://img.shields.io/badge/arXiv-2605.22344-b31b1b.svg)](https://arxiv.org/abs/2605.22344)16[![Project Page](https://img.shields.io/badge/Project-Page-blue.svg)](https://bernini-ai.github.io/)17[![HuggingFace](https://img.shields.io/badge/%F0%9F%A4%97%20HuggingFace-Models-yellow)](https://huggingface.co/collections/ByteDance/bernini)18 19</div>20 21## 🎉 News22 23- **[2026-06-09]** We open-sourced the **1.3B** weights of the Bernini Renderer (**Bernini-R**) on [ByteDance/Bernini-R-1.3B-Diffusers](https://huggingface.co/ByteDance/Bernini-R-1.3B-Diffusers). Fine-tuned from Wan2.1-1.3B, the model performs close to the 14B variant on simple tasks such as style transfer, subtitle or watermark removal, and local editing, while lagging behind on more complex tasks such as human generation.24- **[2026-06-01]** We open-sourced the inference code and model weights of the Bernini Renderer (**Bernini-R**).25- **[2026-05-22]** We released our paper [Bernini: Latent Semantic Planning for Video Diffusion](https://arxiv.org/abs/2605.22344).26 27## ✨ Highlights28 29Bernini is a unified framework for video generation and editing that combines an MLLM-based semantic planner with a DiT-based renderer.30 31On video editing, Bernini reaches the first tier among leading closed-source32commercial models. The leaderboard below comes from our self-built arena33platform, where human annotators blindly vote on paired edits and the votes are34aggregated into a Bradley-Terry score and a pairwise win-rate matrix.35 36<img src="assets/arena.png" width="900" alt="Video editing arena: Bradley-Terry leaderboard and pairwise win-rate matrix"/>37 38## 📦 Installation39 40### Requirements41 42- **Python** 3.11.2.43- **CUDA GPU** — a Hopper GPU (H100/H800/H200) is recommended so FlashAttention-344  can be used; other CUDA GPUs fall back to FlashAttention-2 or PyTorch SDPA.45- **CUDA toolkit** 12.4 (matches the pinned `torch==2.5.1+cu124`; 12.3+ is the46  minimum if you build FlashAttention-3).47- Pinned in `requirements.txt`: `torch==2.5.1+cu124`, `diffusers==0.35.2`,48  `accelerate==0.34.2`, `transformers==4.57.3`.49 50Reference environment (Bernini-R is developed and tested on this setup):51 52| Component | Version      |53|-----------|--------------|54| GPU       | NVIDIA H100  |55| CUDA      | 12.4         |56| Python    | 3.11.2       |57| PyTorch   | 2.5.1+cu124  |58 59### Install60 61```bash62git clone https://github.com/bytedance/Bernini.git bernini && cd bernini63pip install -r requirements.txt64```65 66Optional extras:67 68- **Multi-GPU sequence parallel** needs [Open-VeOmni](https://github.com/ByteDance-Seed/VeOmni)69  (Apache-2.0, Python 3.11). Use `--no-deps` so VeOmni does not pull in a70  different torch build and override the pinned `torch==2.5.1+cu124`:71  `pip install --no-deps git+https://github.com/ByteDance-Seed/VeOmni.git@v0.1.10`.72  Single-GPU inference does not need it.73- **Faster attention** (auto-detected if installed; otherwise PyTorch SDPA is used):74  - FlashAttention-2 — general CUDA GPUs (incl. A100/A800): `pip install flash-attn==2.8.3`.75  - FlashAttention-3 — Hopper only (H100/H800/H200, CUDA ≥ 12.3, PyTorch ≥ 2.4).76    `flash_attn_interface` is not on PyPI; build it from the77    [flash-attention](https://github.com/Dao-AILab/flash-attention) repo's78    `hopper/` directory at tag `v2.8.3`:79    ```bash80    git clone https://github.com/Dao-AILab/flash-attention.git81    cd flash-attention && git checkout v2.8.382    cd hopper && MAX_JOBS=$(nproc) python3 setup.py install --user83    ```84 85### Weights86 87#### Model Performance88 89| Model | EditVerse | OpenVE | OpenS2V | VBench | Bernini-v2v (OS) | Bernini-vr2v (OS) |90|---|---|---|---|---|---|---|91| [Bernini-R 1.3B](https://huggingface.co/ByteDance/Bernini-R-1.3B-Diffusers) | 7.74 | 3.65 | 62.18 | 84.69 | 3.15 | 3.21 |92| [Bernini-R 14B](https://huggingface.co/ByteDance/Bernini-R-Diffusers) | 7.99 | 3.78 | 62.94 | 84.64 | 3.25 | 3.34 |93 94Bernini-R provides two ways to obtain the renderer weights. The **diffusers95format is recommended** — it is a self-contained diffusers-format directory whose96`transformer` / `transformer_2` already hold the Bernini-R weights, so you point97`--config` at it and the weights load directly, with **no** `--high_noise_ckpt` /98`--low_noise_ckpt` needed.99 100#### Option A — diffusers format (recommended)101 102A single ready-to-use diffusers-format model from103[`ByteDance/Bernini-R-Diffusers`](https://huggingface.co/ByteDance/Bernini-R-Diffusers).104It bundles the Wan2.2 base components (VAE, UMT5 text encoder, tokenizer) together105with the Bernini-R transformer weights, so nothing else is downloaded at runtime.106 107```bash108pip install -U "huggingface_hub"109hf download ByteDance/Bernini-R-Diffusers --local-dir Bernini-R-Diffusers110```111 112Then pass it via `--config` and omit the checkpoint flags, e.g.:113 114```bash115python infer_single_gpu.py --config Bernini-R-Diffusers \116    --case assets/testcases/t2i/t2i.json --num_frames 1117```118 119#### Option B — separate checkpoints120 121The original layout, where Bernini-R uses two sets of weights loaded separately:122 1231. **Wan2.2 base** — [`Wan-AI/Wan2.2-T2V-A14B-Diffusers`](https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B-Diffusers) on Hugging Face. Supplies the124   VAE, UMT5 text encoder, tokenizer, and the transformer architecture/base weights.125   It is downloaded automatically on first run (configured by `wan22_base` in126   `configs/bernini_renderer_wan22/config.json`).1272. **Bernini-R checkpoint** — the trained high-noise / low-noise transformer weights128   (safetensors) from [ByteDance/Bernini-R](https://huggingface.co/ByteDance/Bernini-R), passed with129   `--high_noise_ckpt` / `--low_noise_ckpt`. Both a local directory and a Hugging130   Face repo id are accepted.131 132Download models using huggingface-cli:133 134```bash135pip install -U "huggingface_hub"136hf download Wan-AI/Wan2.2-T2V-A14B-Diffusers --local-dir Wan2.2-T2V-A14B-Diffusers137hf download ByteDance/Bernini-R --local-dir Bernini-R138```139 140## 🚀 Usage141 142A run is described by a **case file** — a small JSON under143[`assets/testcases/`](assets/testcases/) that bundles one task's routing and144inputs (`task_type`, `guidance_mode`, `prompt`, source media, `output`). This145keeps long prompts out of the command line. Each task has a directory under146`assets/testcases/` holding one or more case files; see147[`assets/testcases/`](assets/testcases/) for the format and the bundled148`t2i` / `i2i` / `t2v` / `v2v` / `rv2v` /`r2v` examples.149 150### Prompt enhancer (highly recommended)151 152`--use_pe` enhances the prompt through an OpenAI-compatible endpoint and is153recommended for best generation quality. The `openai` SDK is installed by154`requirements.txt`; configure the endpoint with environment variables:155 156```bash157export BERNINI_PE_API_KEY=...      # or OPENAI_API_KEY158export BERNINI_PE_BASE_URL=...     # or OPENAI_BASE_URL159export BERNINI_PE_MODEL=...        # vision-capable chat model160```161 162### Examples by task type163 164Unless an example specifies otherwise, inference outputs **480p / 16fps** (the165defaults — `--max_image_size 848`, `--fps 16`).166 167Each example runs a bundled case in168[`assets/testcases/`](assets/testcases/) — replace `<hi>` / `<lo>` with your169high-/low-noise checkpoint paths. The image tasks (`t2i`, `i2i`) are shown on a170single GPU; the video tasks on 8 GPUs via `torchrun`, where `--ulysses N` gives171N-way Ulysses sequence parallel per sample and the remaining `world_size / N`172ranks run data parallel over the task list. The two scripts take the same173inputs, so any example can be run either way.174 175Inputs can also be passed directly as flags instead of `--case` (`--prompt`,176`--task_type`, `--guidance_mode`, `--video`, `--image`, `--images`,177`--output`); generation parameters (`--seed`, `--num_frames`, ...) are always178command-line flags.179 180**Text-to-image** (`t2i`) — single GPU; generates one frame, so pass `--num_frames 1`181 182```bash183python infer_single_gpu.py --high_noise_ckpt <hi> --low_noise_ckpt <lo> \184    --case assets/testcases/t2i/t2i.json --num_frames 1185```186 187**Image editing** (`i2i`) — single GPU; generates one frame, so pass `--num_frames 1`188 189```bash190python infer_single_gpu.py --high_noise_ckpt <hi> --low_noise_ckpt <lo> \191    --case assets/testcases/i2i/i2i.json --num_frames 1192```193 194**Text-to-video** (`t2v`)195 196```bash197torchrun --nproc-per-node 8 infer_multi_gpu.py \198    --high_noise_ckpt <hi> --low_noise_ckpt <lo> --ulysses 8 \199    --case assets/testcases/t2v/t2v.json200```201 202**Video editing** (`v2v` / `mv2v`) — two cases are provided.203 204For edits where the main subject keeps its ordinary motion (case 1 adds a205snowman to the scene), the `v2v` task type is enough:206 207```bash208torchrun --nproc-per-node 8 infer_multi_gpu.py \209    --high_noise_ckpt <hi> --low_noise_ckpt <lo> --ulysses 8 \210    --case assets/testcases/v2v/v2v_case1.json211```212 213For edits that need to change the subject's motion (case 2 makes the person214crouch down), the `mv2v` task type gives better results:215 216```bash217torchrun --nproc-per-node 8 infer_multi_gpu.py \218    --high_noise_ckpt <hi> --low_noise_ckpt <lo> --ulysses 8 \219    --case assets/testcases/v2v/v2v_case2.json220```221 222**Reference + video editing** (`rv2v`) — two cases are provided.223 224Case 1 is reference-image-guided video editing — replacing a garment in the225source video with one from a reference image:226 227```bash228torchrun --nproc-per-node 8 infer_multi_gpu.py \229    --high_noise_ckpt <hi> --low_noise_ckpt <lo> --ulysses 8 \230    --case assets/testcases/rv2v/rv2v_case1.json231```232 233Case 2 is a video-insertion example — inserting content into the source video.234It is run at 720p / 24fps to show the insertion result more clearly:235 236```bash237torchrun --nproc-per-node 8 infer_multi_gpu.py \238    --high_noise_ckpt <hi> --low_noise_ckpt <lo> --ulysses 8 \239    --case assets/testcases/rv2v/rv2v_case2.json \240    --num_frames 121 --fps 24 --max_image_size 1280241```242 243**Reference-to-video** (`r2v`) — drives a video from one or more reference images244 245```bash246torchrun --nproc-per-node 8 infer_multi_gpu.py \247    --high_noise_ckpt <hi> --low_noise_ckpt <lo> --ulysses 8 \248    --case assets/testcases/r2v/r2v.json249```250 251See `python infer_single_gpu.py --help` for the full argument list.252 253### Gradio demo254 255`gradio_demo.py` exposes the same pipeline through a Gradio UI: the task-type256dropdown auto-fills `guidance_mode` (still user-editable), uploaded media is257routed to the matching slot, and the result is rendered inline.258 259```bash260# Single GPU261python gradio_demo.py --high_noise_ckpt <hi> --low_noise_ckpt <lo> --port 7860262 263# 8 GPUs, 8-way Ulysses sequence parallel264torchrun --nproc-per-node 8 gradio_demo.py --ulysses 8 \265    --high_noise_ckpt <hi> --low_noise_ckpt <lo> --port 7860 --share266```267 268Add `--use_pe` (and `export OPENAI_API_KEY=...` / `BERNINI_PE_API_KEY=...`) to269enable GPT prompt enhancement; the in-UI checkbox is a per-request switch on270top of this flag.271 272## 📑 Citation273 274If you use Bernini in your research, please cite:275 276```bibtex277@article{bernini,278  title   = {Bernini: Latent Semantic Planning for Video Diffusion},279  author  = {Chenchen Liu and Junyi Chen and Lei Li and Lu Chi and Mingzhen Sun and Zhuoying Li and Yi Fu and Ruoyu Guo and Yiheng Wu and Ge Bai and Zehuan Yuan},280  journal = {arXiv preprint arXiv:2605.22344},281  year    = {2026}282}283```284 285## 🙏 Acknowledgements286 287Bernini builds on several outstanding open-source projects:288 289- [Wan2.2-T2V-A14B](https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B)290- [Qwen2.5-VL-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-VL-7B-Instruct)291- [VeOmni](https://github.com/ByteDance-Seed/VeOmni)292 293We thank the authors and communities of these projects for their contributions.294 295## 📄 License296 297Apache License 2.0. See [LICENSE](LICENSE).