nvidia/Diffusion_Renderer_Forward_Cosmos_7B
7299
1---2license: other3license_name: nvidia-open-model-license4license_link: >-5 https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/6---7 8# DiffusionRenderer (Cosmos): Neural Inverse and Forward Rendering with Video Diffusion Models9 10[Ruofan Liang](https://www.cs.toronto.edu/~ruofan/)\*, [Zan Gojcic](https://zgojcic.github.io/), [Huan Ling](https://www.cs.toronto.edu/~linghuan/), [Jacob Munkberg](https://research.nvidia.com/person/jacob-munkberg), [Jon Hasselgren](https://research.nvidia.com/person/jon-hasselgren), [Zhi-Hao Lin](https://chih-hao-lin.github.io/), [Jun Gao](https://www.cs.toronto.edu/~jungao/), [Alexander Keller](https://research.nvidia.com/person/alex-keller), [Nandita Vijaykumar](https://www.cs.toronto.edu/~nandita/), [Sanja Fidler](https://www.cs.toronto.edu/~fidler/), [Zian Wang](https://www.cs.toronto.edu/~zianwang/)\*11 12\* indicates equal contribution13 14**[Paper](https://arxiv.org/abs/2501.18590) | [Project Page](https://research.nvidia.com/labs/toronto-ai/DiffusionRenderer/)**15 16 17## Description: 18Cosmos Diffusion Renderer is a video diffusion model that accurately estimates geometry and material buffers, and generates photorealistic images under specified lighting conditions, offering fundamental tools for image editing applications. This model is available for commercial use. 19 20 21### License/Terms of Use:22The model is distributed under [NVIDIA Open Model License](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/). 23 24### Deployment Geography:25Global 26 27### Use Case:28AI research, development and benchmarking for image/video delighting and relighting tasks. 29 30### Release Date:31Github 06/11/2025 via https://github.com/nv-tlabs/cosmos1-diffusion-renderer32 33## Reference(s):34Project page: https://research.nvidia.com/labs/toronto-ai/DiffusionRenderer/35 36## Model Architecture:37**Architecture Type:** Transformer.38 39**Network Architecture:** Transformer. This model has 7B model parameters.40 41## Input: 42**Input Type(s):** Image, Video43 44**Input Format(s):** Red, Green, Blue (RGB); Video input are frames of images. 45 46**Input Parameters:**47The input video data are five-dimensional, with the input dimension specified as [batch_size, num_frames, height, width, 3], where the three channels are Red, Green, Blue (RGB) channels. 48The input image data are five-dimensional [batch_size, 1, height, width, 3], with the frame channel set to 1. 49 50**Other Properties Related to Input:** The input resolution is 704 x 1280.51 52## Output: 53**Output Type(s):** Image, Video54 55**Output Format(s):** Red, Green, Blue (RGB); Output video is saved as image frames. 56 57**Output Parameters:**58The output video data are five-dimensional, with the dimension specified as [batch_size, num_frames, height, width, 3], where the three channels are Red, Green, Blue (RGB) channels. 59The output image data are five-dimensional [batch_size, 1, height, width, 3], with the frame channel set to 1. 60 61**Other Properties Related to Output:** The input resolution is 704 x 1280.62 63Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA’s hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.64 65 66## Software Integration:67 68**Runtime Engine(s):**69Not Applicable - Uses Python Scripts and Pytorch70 71**Supported Hardware Microarchitecture Compatibility:**72*NVIDIA Ampere (A100)*73 74 75**[Preferred/Supported] Operating System(s):**76['Linux']77 78**Model Version(s):**79- `Diffusion_Renderer_Inverse_Cosmos_7B`: estimates geometry and material buffers from input image or video. 80- `Diffusion_Renderer_Forward_Cosmos_7B`: generates photorealistic images or videos from G-buffers and environment map. 81 82 83[//]: # (# Training Datasets:)84 85## Training Dataset:86 87The training data contains samples with paired data of RGB video frames, G-buffers (base color, roughness, metallic, normals, depth), and lighting (HDR environment maps). 88 89**Data Collection Method**: Synthetic90 91**Labeling Method**: Hybrid: Synthetic, Automated 92 93**Properties:** 94- 150k synthetic videos, 57 frames each at 704×1280 resolution.95- 150k auto-labeled real-videos, 57 frames each at 704×1280 resolution.96 97For additional details, please refer to the [paper](https://research.nvidia.com/labs/toronto-ai/DiffusionRenderer/). 98 99 100## Inference:101 102**Engine:** Tensor(RT)103 104**Test Hardware:**105 106A100 GPUs107 108## Ethical Considerations:109NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse. 110 111Please report security vulnerabilities or NVIDIA AI Concerns [here](https://www.nvidia.com/en-us/support/submit-security-vulnerability/).112 