Favicon of Mochi 1

Mochi 1

An open-source text-to-video model you can run locally through ComfyUI or Python, with Apache 2.0 licensing and LoRA fine-tuning.

Screenshot of Mochi 1 website

Mochi 1 is a text-to-video model for creators and developers who want to generate videos on their own hardware or adapt a model to their own footage. Genmo releases it under Apache 2.0, with downloadable weights and code for local use. Genmo also offers a hosted playground for trying the model in a browser.

The model focuses on photorealistic video, with motion quality and adherence to written prompts as central goals. It generates at 480p and is less suited to animated styles. Extreme motion can produce warping or other distortions, a limitation that matters for action scenes and fast camera movement.

You can use Mochi through ComfyUI or its Python tools, which include a Gradio interface and an API for incorporating generation into other software. LoRA training lets you adapt it using your own videos, and it accepts fine-tuned LoRA files in safetensors format. Its Asymmetric Diffusion Transformer architecture gives more capacity to visual processing than text processing to reduce inference memory requirements.

Hardware needs depend on the implementation. The repository supports splitting the model across multiple GPUs, while its single-GPU path needs roughly 60GB of VRAM; Genmo recommends an H100. ComfyUI optimizations can bring inference below 20GB of VRAM. Fine-tuning requires one H100 or A100 GPU with 80GB of memory.

Similar to Mochi 1