Skip to main content
HappyHorse 1.1 is the latest release of Alibaba’s production-grade video generation model, now available in ComfyUI as a Partner Node. It is engineered for real-world creative production: short episodic series, e-commerce commercials, brand marketing content, and game cutscenes. A standout feature is native synchronized audio generation: HappyHorse 1.1 produces dialogue, sound effects, and background music in a single render pass, with audio tightly bound to the visual timeline. There are no extra audio steps. Version 1.1 targets five core production-critical capabilities: dynamic expressive motion, consistent character rendering, reliable prompt adherence, stable text rendering, and authentic cinematic framing.

What HappyHorse 1.1 is good at

  • Native audio-video sync: Dialogue, SFX, and background music in one pass (no extra steps)
  • Three creation paths: Text-to-Video (T2V), Image-to-Video (I2V), Reference-to-Video (R2V)
  • Multi-image R2V: Up to 9 reference images per generation with preserved identity
  • Multi-character consistency: Multiple character references stay distinct with no visual cross-contamination
  • Character × scene separation: Feed characters and scenes as independent references; characters stay consistent as the background changes
  • Long-context prompts: Handles prompts beyond 2,500 characters; a single prompt can describe 6–8 consecutive scenes with the model autonomously allocating time and switching camera angles
  • Cinematic language: Full support for shot-reverse-shot, tracking shot, and cohesive transitions/pacing between shots
  • Flexible output: 720p and 1080p, 3–15 seconds, aspect ratios 16:9 / 9:16 / 1:1 / 4:3 / 3:4 / 21:9 and more
  • Production-ready visuals: Fixes shiny skin and over-sharpening issues from v1.0 for natural-looking close-ups and series content

Use it in ComfyUI

HappyHorse 1.1 workflows

Run the text-to-video, image-to-video, and reference-to-video workflows in ComfyUI, locally or on Comfy Cloud