text-to-video

1 posts

netflix

Toward More Controllable AI Video Editing: An Early Research Exploration at Netflix (opens in new tab)

Netflix explores AI video-editing tools designed to preserve artists’ creative control rather than regenerate entire clips indiscriminately. The research addresses two major problems: unintended changes to untouched footage and physically implausible results when objects are removed. Its proposed systems, Vera and VOID, generate targeted edits while preserving scene identity, performance, and continuity. ## Challenges in Generative Video Editing - Full-video regeneration can unintentionally change: - Actors’ identities and performances - Backgrounds and objects - Important scene details - Object removal often produces unnatural results because models erase the target without reconstructing realistic motion and physical interactions. - Professional editors need precise control over what changes and what remains untouched. ## Vera: Layered Video Diffusion - Vera generates: - An edit layer containing the requested visual change - An alpha matte defining where that change should appear - These layers are composited with the original footage, leaving pixels outside the edited region intact. - The approach supports tasks such as: - Adding objects - Changing backgrounds - This layered design helps preserve original identities, performances, and details. ## Training Dataset - Netflix created a custom dataset because existing public datasets lacked high-quality layered video data. - The dataset contains 486,000 frames at 832×480 resolution. - It includes: - **Synthetic composites:** Foreground objects with alpha mattes placed over generated backgrounds. - **Realistic single-object videos:** Real footage processed with segmentation, matting, background generation, and human review. - **Realistic multi-object videos with effects:** Objects isolated along with shadows, reflections, and other scene effects. ## Vera’s Model Architecture - Vera uses a Mixture-of-Transformers design with three specialized DiTs for: - The edit layer - The alpha matte - The composite video - Each branch has its own attention projections and feed-forward weights, allowing specialization while joint attention enables communication between layers. - The model is initialized from a pretrained text-to-video model. - Additional embeddings and input layers help distinguish source-video, mask, alpha, and composite information. ## Evaluation and Results - Netflix tested Vera on: - 72 object-addition video-prompt pairs - 69 background-change pairs - The benchmark included varied motion speeds, camera movements, object counts, and scene complexity. - Evaluation measured: - Preservation of untouched content - Compliance with text instructions - Temporal and per-frame video quality - Vera-1.3B and Vera-14B substantially outperformed existing methods on content preservation while achieving comparable instruction-following and visual quality. Netflix’s research favors localized, layered editing over unrestricted video regeneration. Vera demonstrates how separating edits from original footage can make generative tools safer and more controllable for professional workflows; the accompanying VOID research aims to apply similar principles to physically plausible object and interaction removal.