Open-source video generation has moved quickly since the first research systems. In 2026, the most useful projects cover text-to-video, image-to-video, editing, audio-video generation, and local workflows, but they differ sharply in quality, speed, hardware requirements, and license terms. If you need stills rather than clips, compare the best AI image generators in 2026 instead.
Here are nine open-source video generation model families worth knowing in 2026:
1. Wan2.2
Alibaba's current downloadable open-weight flagship, released under Apache 2.0. The suite pairs a 27B Mixture-of-Experts design—separate high-noise and low-noise experts for layout and detail—with a dense TI2V-5B model that handles both text-to-video and image-to-video at 720p and 24 fps on a single consumer GPU, thanks to a high-compression VAE. Later Wan-branded hosted releases do not have downloadable model repositories in the official Wan-Video organization, so Wan2.2 remains its newest downloadable video-model release. Official repository
2. Wan2.1
Still worth knowing as the low-VRAM entry point. The 1.3B text-to-video model needs about 8.19 GB of VRAM for a five-second 480p generation, and the broader suite covers text-to-video, image-to-video, video editing, text-to-image, and video-to-audio workflows. Official repository
3. HunyuanVideo 1.5
Tencent's November 2025 rebuild dropped from 13B to 8.3B parameters while improving quality, and it now runs on consumer hardware—roughly 14 GB of VRAM minimum with offloading and comfortably on a 24 GB card. A dedicated super-resolution network pushes output to 1080p. Note the custom community license: it does not apply in the European Union, United Kingdom, or South Korea. Official repository
4. LTX-2 → LTX-2.5
Lightricks released LTX-2's weights, inference package, and training tools in January 2026, and now recommends LTX-2.5, superseding LTX-2 and LTX-2.3 in its official setup instructions. The 22B model generates synchronized audio and video in one pass, with full and distilled checkpoints and workflows for text-to-video, image-to-video, audio-conditioned generation, and editing. Its multishot generation can produce connected scenes with continuity across cuts. High-resolution workflows use refinement and upscaling.
The LTX-2.x Community License permits qualifying noncommercial use and commercial use by entities below $10 million in annual revenue; larger organizations generally need a paid license for commercial use. Distilled checkpoints offer faster local generation. Official repository | Official model card LTX-2.5
5. MiniMax H3
MiniMax’s audio-video generation family uses a 33B-parameter transformer to produce 4–15-second clips at 24 fps with native stereo sound. Two downloadable checkpoints cover text and first/last-frame conditioning, plus generation guided by reference images, videos, and audio. Local H3-Base runs at 768p; the full 2K workflow also requires hosted context-processing and regeneration modules that are not included in the open release. The weights use the custom MiniMax H3 Community License. Official repository
6. CogVideoX
The open CogVideo family includes text-to-video and image-to-video checkpoints in several sizes. Smaller variants can run on consumer GPUs with memory-saving settings, making the family useful for experimentation. CogVideoX1.5-5B and CogVideoX1.5-5B-I2V as its latest checkpoints. They support clips up to ten seconds, with 1360 × 768 text-to-video output and flexible image-to-video dimensions within documented limits. CogVideoX’s licensing differs by checkpoint: CogVideoX-2B is Apache 2.0, while some 5B releases use the custom CogVideoX license. Official repository
7. Mochi 1
Genmo's 10B-parameter Asymmetric Diffusion Transformer is released with code and weights under Apache 2.0. The reference implementation is hardware-heavy, although community runtimes can reduce memory requirements substantially. Official repository
8. Kandinsky 5.0
Kandinsky Lab's family spans a 2B Video Lite line with checkpoints for clips up to 10 seconds and a 19B Video Pro line. The official project reports that Video Pro ranked first among open-source text-to-video models on LMArena in December 2025, and the family offers strong English and Russian prompt understanding. The repository is MIT licensed, but always confirm the license on the exact checkpoint. Official repository
9. Stable Video Diffusion
Stability AI’s SVD XT 1.1 turns still images into 25-frame videos at 1024 × 576, with fine-tuning at six fps to improve output consistency. Although older than most models on this list, its Diffusers integration makes it a practical baseline for research, comparisons, and prototypes. The weights use a custom license; check Stability AI’s terms before commercial use. Official model card
Four low-friction ways to try these models:
Use an official Hugging Face demo or model Space when one is available.
Install a maintained ComfyUI workflow for the selected checkpoint.
Use the project's official local interface or command-line example.
Test through a hosted inference provider before investing in local GPUs.
Trying a checkpoint and running one are different problems. For the serving side, we list open-source tools for model deployment, from local models to production inference.
FAQ
What is the best open-source AI video generator in 2026?
There is no single best model for every workflow. Wan2.2 is the strongest all-round downloadable open-weight starting point, HunyuanVideo 1.5 offers a strong quality-to-hardware ratio, LTX-2 adds native synchronized audio and 4K generation, and Kandinsky 5.0 Lite and CogVideoX offer accessible model sizes. Test the same prompts across two or three candidates before choosing.
Can I run an open-source video generation model locally?
Yes, and the hardware bar has dropped considerably. Wan2.1's 1.3B and Wan2.2's TI2V-5B models target consumer-grade GPUs, and HunyuanVideo 1.5 can run with 14 GB of VRAM when model offloading is enabled. Larger checkpoints, such as Wan2.2's 27B MoE models or the reference Mochi 1 pipeline, still benefit from high-memory GPUs, quantization, or a hosted service.
How much VRAM do open-source video models need?
Requirements range from roughly 8 GB for optimized small checkpoints to 60 GB or more for large reference implementations. Resolution, frame count, precision, attention implementation, and CPU offloading can change the practical requirement, so always check the specific model card and runtime documentation.
Are open-source AI video generators free for commercial use?
Not automatically. Wan2.2 and Mochi 1 use Apache 2.0, while Kandinsky 5.0's current repository is MIT licensed. CogVideoX licensing varies by checkpoint; LTX-2 requires a paid license for entities with at least $10 million in annual revenue; HunyuanVideo 1.5 excludes the EU, UK, and South Korea; and Stable Video Diffusion uses Stability AI's community terms. Review the exact checkpoint license before using generated video in a product or client project.
What is the difference between text-to-video and image-to-video?
Text-to-video generates a sequence from a written prompt, so the model controls both appearance and motion. Image-to-video starts from a supplied frame and adds motion, which usually gives the creator more control over composition, characters, and visual identity.





