Open Source and Self Hosted Video Models

Weights you can download, run on your own graphics card and fine-tune, plus the software that makes that practical. Teams take this route for three reasons: no per-clip credit meter, footage that never leaves their own machines, and control over the model itself. The trade is setup time and hardware. This page covers the open model families and the node editors, launchers and hosted runners people actually use to drive them. We rank on output next to the hosted rivals, licence terms for commercial work, the video memory you realistically need, and how active the project is, because an unmaintained repository ages badly in this field.

  1. 1

    Alibaba's open large-scale video foundation model suite supporting text-to-video, image-to-video, video editing, and more, with 1.3B and 14B variants runnable on consumer GPUs and uniquely capable of rendering both Chinese and English text in generated video.

  2. 2

    Professional node-based workflow engine for AI-driven image, video, and audio generation with 60,000+ community nodes, extensive shared workflows, and enterprise adoption by Amazon Studios, Netflix, Apple, and others. Runs open models locally with full parameter visibility.

  3. 3

    Tencent's largest open-source video model with 13+ billion parameters, using a multimodal LLM text encoder and 3D causal VAE for text-to-video and image-to-video generation at high resolution.

  4. 4

    An 11-billion-parameter open-source video model supporting text-to-video and image-to-video generation at 256p and 768p resolutions, with flexible frame counts up to 129 frames. Apache 2.0 licensed.

  5. 5

    Stability AI's open image-to-video diffusion model that generates short video clips from still images, available for research and commercial licensing.

  6. 6

    Zhipu AI's open-source video generation framework supporting text-to-video, image-to-video, and video continuation with multiple model sizes and resolutions.

  7. 7

    Alibaba PAI's open transformer-based video generation framework with 7B and 12B model variants, supporting text-to-video, image-to-video, video-to-video, and multiple control modes (pose, depth, Canny edge, etc.), deployable via ComfyUI, web interface, or Python with flexible VRAM tiers.

  8. 8

    Lightricks' open-weight video model with 13B and 2B variants, real-time generation capability, and support for text-to-video, image-to-video, and video extension under Apache-2.0 and OpenRail-M licenses.

  9. 9

    A MIT-licensed video model using autoregressive flow matching, generating 10-second 768p videos at 24 FPS with text-to-video and image-to-video support. Includes single-GPU and multi-GPU inference plus CPU offloading options.

  10. 10

    Genmo's open-source text-to-video model that converts written descriptions into video, with weights and code available for local deployment and customization.

  11. 11

    Open-source motion module that converts Stable Diffusion text-to-image models into animation generators without retraining, with support for SD1.5 and SDXL, plus SparseCtrl and MotionLoRA extensions.

  12. 12

    Next-frame prediction neural network using a 13B model that generates long-form video on consumer GPUs (6GB VRAM minimum), features a Windows one-click installer, and generates with workload invariant to video length via constant-input-context architecture.