Happy Horse 1.1: What's New, Specs, and Benchmarks

Happy Horse 1.1 (June 2026): improved motion, temporal consistency, and visual fidelity on the Transfusion architecture. Specs, benchmarks, and upgrade guide.

By VioEvo EditorialPublished July 20, 2026Reading time 8 min

Before reading, try Happy Horse 1.1.

Released: June 23, 2026 · Developer: Alibaba Taotian Group · Future Life Lab · Status: Current production version


Version note: This page covers Happy Horse 1.1 specifically — the refinement release from June 2026. For the model family overview, architecture deep dive, and version comparison, see the Happy Horse family guide. For the original release history and launch benchmarks, see Happy Horse 1.0.


What Changed in 1.1

Happy Horse 1.1 was released on June 23, 2026 — ten weeks after 1.0's anonymous leaderboard debut. It is a refinement release on the same Transfusion architecture: the parameter count, layer structure, generation speed, and output specifications are unchanged. The update targets the specific weaknesses most frequently reported in 1.0 production use.

Motion Expressiveness

The most visible improvement. Version 1.0 occasionally produced motion that felt sluggish or floating at complex or high-speed actions — running, fighting, rapid camera tracking, sequences with multiple interacting characters. Version 1.1 rebuilds the motion modeling to address this directly:

  • High-speed actions carry more physical momentum and feel grounded rather than floating
  • Camera tracking maintains consistent motion continuity through acceleration and deceleration
  • Character interactions have better dynamic contact: collisions, handoffs, and physical exchanges feel more realistic
  • Fine motor movements (hands, facial expressions, secondary body motion) resolve more naturally

Temporal Consistency

Subject drift (where faces, logos, costumes, or objects gradually shift in appearance during a clip) was the second most commonly reported limitation of 1.0, particularly in longer sequences or sequences with complex backgrounds.

Version 1.1 addresses this across three dimensions:

  • Character identity holds more reliably through movement, lighting changes, and camera angle changes within a single clip
  • Scene elements (backgrounds, props, lighting setups) remain stable through the duration of the generation
  • Style registers: if a photorealistic or cinematic style is specified, it does not drift toward a less defined look as the clip progresses

The improvement is most pronounced in 10–15 second clips, where drift was most visible in 1.0.

Visual Fidelity

Sharper facial textures and improved fine detail, particularly visible in:

  • Close-up portrait and speaking-head shots
  • Complex fabric and material rendering
  • High-detail environments with multiple textured surfaces

The improvement is incremental rather than transformative. Happy Horse 1.0 was already strong on visual quality; it is consistently measurable in close-up comparisons.

Prompt Adherence

Version 1.1 follows complex multi-element prompts more reliably. Specifically:

  • Camera language specifications (push-in speed, rack focus timing, handheld drift intensity) are rendered with greater precision
  • Lighting direction and quality specified in prompts carry through more consistently
  • Multi-element compositions where subject, action, environment, camera, and atmosphere are all specified produce results that honor the full specification rather than partially collapsing to dominant elements

Version 1.1 Specifications

SpecificationDetails
Parameters15 billion
ArchitectureTransfusion (Unified Multimodal: Autoregressive + Diffusion)
Layers40 (4 modality-specific input + 32 shared + 4 modality-specific output)
Denoising steps8 (DMD-2 distillation, no CFG required)
Clip duration3–15 seconds
Native resolution720p / 1080p
Aspect ratios16:9 · 9:16 · 4:3 · 3:4 · 1:1
Frame rate24 FPS
Generation speed~38 seconds (single NVIDIA H100)
Audio outputNative joint: dialogue · SFX · ambient · BGM
Lip sync languages7 (EN · ZH-Hans · ZH-Hant · JA · KO · DE · FR)
Lip sync WER14.60%
WorkflowsT2V · I2V · R2V
Reference imagesUp to 9
LicenseApache 2.0

Source: Artificial Analysis Video Arena · Alibaba Taotian Group · as of July 2026


Benchmark Performance

Happy Horse 1.1's Elo scores on Artificial Analysis Video Arena as of July 2026:

CategoryHappy Horse 1.1 Elo
Text-to-Video (no audio)1,273
Text-to-Video (with audio)1,150
Image-to-Video (with audio)1,110

Source: Artificial Analysis Video Arena · July 2026 · Elo scores are live and shift as new votes are collected

Reading These Numbers

The 1.1 Elo scores are lower than the 1.0 peak scores from April 2026. This does not indicate a regression between versions. Two factors explain the gap:

The competitive field changed. Happy Horse 1.0 debuted into a leaderboard where the strongest competitor was Seedance 2.0 at approximately 1,270 T2V no-audio Elo. Between April and July 2026, multiple new models entered the arena, raising the competitive baseline and redistributing Elo across a larger field. The same pattern is observable for every model that was #1 in April: their Elo scores reflect a differently populated leaderboard.

The with-audio gap narrowed. In the with-audio T2V category, Happy Horse 1.0 trailed Seedance 2.0 by approximately 14 points at launch. Version 1.1 reduced this gap. The improved audio-visual synchronization and tighter prompt adherence both contribute to better preference voting in audio-inclusive blind tests.


Comparing 1.0 and 1.1

The following table covers the specific dimensions where the two versions differ. For specifications that are identical, see the family overview comparison table.

Dimension1.01.1
High-speed motionOccasionally sluggishRebuilt: more physical momentum
Subject driftVisible in 10–15s clipsSignificantly reduced
Facial detailStrongImproved: sharper textures
Complex prompt followingStrongImproved: more reliable
With-audio leaderboard~1,205 Elo (April 2026)1,150 Elo (July 2026, larger field)
Video editing workflowSupportedNot listed as a supported workflow in 1.1

Note on the video editing workflow: Happy Horse 1.0 included a video-edit endpoint. Whether this workflow is available via 1.1 varies by platform and API provider. Confirm availability with your specific provider before building a 1.1-based editing pipeline.


Real-World Use Cases for 1.1

Action and movement content

The rebuilt motion model in 1.1 makes it the better choice for any content where physical believability of movement is primary: sports sequences, dance, action choreography, running or chasing scenes. The floating quality observed in 1.0 at these scenarios is substantially reduced.

Brand and character consistency work

The reduced subject drift in 1.1 is directly valuable for e-commerce product video, character-consistent narrative content, and any workflow where a specific identity (face, product, logo) must hold stable across the clip duration. Up to nine reference images can be used to lock character or style identity.

Speaking-head and multilingual content

Happy Horse 1.1 maintains the 14.60% WER lip sync across seven languages established in 1.0, with the improved visual fidelity making close-up dialogue sequences more polished. This remains one of the strongest use cases for the Happy Horse family: synchronized dialogue without post-processing, across seven languages including Cantonese as a distinct target.

High-specificity prompt work

Creators working with detailed cinematographic specifications (specific camera language, lighting direction, atmosphere) will find 1.1's improved instruction following more reliable than 1.0. Complex multi-element prompts are more likely to produce output that honors the full specification.


Limitations That Carry Forward from 1.0

15-second maximum clip duration

No native Scene Extension or Video Extend capability. Longer content still requires chaining individual generations.

1080p maximum resolution

4K output is not available. For broadcast-grade 4K requirements, Kling 3.0 and Veo 3.1 remain the current options.

With-audio leaderboard position

While the gap narrowed from 1.0, Happy Horse 1.1 does not hold the top position in the with-audio T2V category as of July 2026. For workflows where audio generation quality is the primary evaluation criterion, verify current leaderboard rankings before selecting a model.


Frequently Asked Questions

What specifically changed from 1.0 to 1.1?

Three primary areas: motion expressiveness (rebuilt motion model, better physical grounding for complex actions), temporal consistency (reduced subject drift, more stable scene and style across clip duration), and visual fidelity (sharper facial textures, better fine detail). Prompt adherence also improved. The architecture, generation speed, and output specifications are unchanged.

Should I upgrade from 1.0 to 1.1?

Yes, for new work. Happy Horse 1.1 addresses the most commonly reported production limitations of 1.0 without any trade-off in speed or output specs. If you have an existing 1.0 pipeline with calibrated prompts, validate 1.1 output before switching, as the motion and consistency changes may affect results in ways that require prompt tuning.

Does 1.1 support the same reference image count as 1.0?

Yes. Up to nine reference images are supported in 1.1, the same as 1.0.

How does 1.1 handle lip sync compared to 1.0?

The core lip sync capability is unchanged: native joint generation at 14.60% WER across seven languages (English, Mandarin, Cantonese, Japanese, Korean, German, French). The visual fidelity improvements in 1.1 make close-up performance shots more polished, which benefits speaking-head content indirectly.

Is 1.1 faster than 1.0?

Generation speed is effectively unchanged — approximately 38 seconds for a 1080p clip on a single NVIDIA H100. The DMD-2 distillation and 8-step denoising that enable this speed are unchanged in 1.1.


For the original launch specs and debut benchmark context: Happy Horse 1.0 · For the family overview: Happy Horse