
Alibaba has launched Wan2.1-VACE (Video All-in-one Creation and Editing), a powerful new open-source model designed to simplify and enhance video creation and editing. This cutting-edge tool combines multiple video processing functions into one unified model, making it easier for users to create and edit videos efficiently.
As part of Alibaba’s Wan2.1 video generation model series, VACE is the first open-source model in the industry to provide a comprehensive solution for a wide range of video generation and editing tasks. It supports multi-modal inputs—including text, images, and video—allowing creators to work seamlessly across various formats. The model’s editing capabilities include referencing specific frames, repainting video content, modifying selected areas, and extending videos in both time and space, offering a flexible and creative toolkit for users.
With Wan2.1-VACE, users can generate videos featuring specific subjects based on sample images, animate static images with natural motion, and apply advanced editing techniques like pose transfer, motion and depth control, and recolorization. The model also enables precise editing of specific regions within a video, such as adding, removing, or modifying elements, without disturbing the surrounding visuals. It can intelligently expand the video’s boundaries, filling in new content to enhance the overall viewing experience.

The all-in-one nature of Wan2.1-VACE allows users to combine multiple functions effortlessly. They can turn static images into videos by controlling object movements, replace characters or items using reference materials, animate selected figures, manipulate poses, or even transform vertical images into horizontal videos by adding contextually appropriate elements.
Wan2.1-VACE is built on innovative technology that considers the diverse needs of video editing during its design. It features a unified interface known as the Video Condition Unit (VCU), which enables streamlined processing of text, image, video, and mask inputs. It also uses a Context Adapter architecture to incorporate task-specific concepts through formalized representations of time and space, giving the model the flexibility to handle a wide array of video synthesis tasks.
These advancements make Wan2.1-VACE highly applicable in a range of scenarios, including social media video production, marketing content creation, film and TV post-production, and educational video development. The model’s design supports rapid, high-quality video generation with significantly reduced costs and time requirements.
To support the broader AI and developer community, Alibaba is releasing Wan2.1-VACE in two versions: a 14-billion-parameter model and a more lightweight 1.3-billion-parameter version. Both are freely available for download on Hugging Face, GitHub, and Alibaba Cloud’s open-source platform, ModelScope.
Alibaba has positioned itself as a leader in AI openness. In February 2025, it released four Wan2.1 models to the public, followed by a video generation model that enables scene creation based on start and end frames. Collectively, these models have been downloaded over 3.3 million times across Hugging Face and ModelScope, highlighting the growing interest in accessible, powerful AI tools for content creation.
One thought on “A powerful open-source tool for video creation and editing”