Happy Horse 1.0: Model Specs, Benchmarks, and Release

Deep dive into Happy Horse 1.0: the model that debuted anonymously at #1 on Artificial Analysis in April 2026. Specs, benchmark data, and capabilities.

By VioEvo EditorialPublished June 9, 2026Updated July 20, 2026Reading time 8 min

Before reading, try Happy Horse 1.1, the current generation.

Released: April 7, 2026 · Developer: Alibaba Taotian Group · Future Life Lab · Status: Available (superseded by 1.1 for most workflows)


Version note: This page covers Happy Horse 1.0 specifically, the initial release that debuted anonymously on the Artificial Analysis leaderboard in April 2026. For the model family overview, architecture deep dive, and version comparison, see the Happy Horse family guide. For the current production version, see Happy Horse 1.1.


The Debut

On April 7, 2026, an anonymous entry appeared on the Artificial Analysis Video Arena leaderboard. No press release. No product page. No identifiable creator. Just a model called "HappyHorse-1.0" that immediately claimed the #1 position in both text-to-video and image-to-video categories, surpassing ByteDance's Seedance 2.0 by nearly 60 Elo points in the no-audio T2V category and setting a new all-time record in image-to-video.

The AI industry spent three days trying to figure out who built it.

On April 10, 2026, Alibaba confirmed the model as its own, built by the Future Life Lab inside Alibaba's Taotian Group (the division that runs Taobao and Tmall), under the newly established ATH AI Innovation Unit. BABA stock jumped more than 4% intraday on the news. Reporting from Bloomberg, CNBC, and The Information confirmed the identity.

The lab is led by Zhang Di, former Vice President at Kuaishou and the technical architect behind Kling AI. That background explains a great deal about what Happy Horse 1.0 is: the person who designed one of the previous generation's strongest video models built this one from scratch, with a full architectural reset.


Version 1.0 Specifications

SpecificationDetails
Parameters15 billion
ArchitectureTransfusion (Unified Multimodal: Autoregressive + Diffusion)
Layers40 (4 modality-specific input + 32 shared + 4 modality-specific output)
Denoising steps8 (DMD-2 distillation, no CFG required)
Clip duration3–15 seconds
Native resolution720p / 1080p
Aspect ratios16:9 · 9:16 · 4:3 · 3:4 · 1:1
Frame rate24 FPS
Generation speed~38 seconds (single NVIDIA H100)
Audio outputNative joint: dialogue · SFX · ambient · BGM
Lip sync languages7 (EN · ZH-Hans · ZH-Hant · JA · KO · DE · FR)
Lip sync WER14.60% (lowest documented, 2026)
Multi-shot consistency~87% (highest documented, 2026)
Visual styles50+
WorkflowsT2V · I2V · R2V · Video Editing
Reference imagesUp to 9
LicenseApache 2.0

Source: Artificial Analysis Video Arena · Alibaba Taotian Group · as of June 2026


Benchmark Performance at Launch

Happy Horse 1.0 debuted on the Artificial Analysis Video Arena on April 7, 2026, immediately claiming top positions across multiple categories. The leaderboard uses blind pairwise comparison: human evaluators see two videos generated from the same prompt and select the better one without knowing which model produced which output.

Artificial Analysis Video Arena, April 2026 (at launch):

CategoryHappy Horse 1.0 EloPosition
Text-to-Video (no audio)1,333–1,357#1
Image-to-Video (no audio)1,391–1,414#1 (all-time record at launch)
Text-to-Video (with audio)~1,205#2 (behind Seedance 2.0 at ~1,219)

Source: Artificial Analysis Video Arena · April 2026 snapshot

The 14-point gap in the with-audio T2V category is meaningful context. Happy Horse 1.0 led Seedance 2.0 by approximately 60 points in no-audio categories, but trailed by 14 points in the with-audio category — reflecting a genuine difference in audio quality at launch rather than overall video quality. For most production use cases the gap was not decisive, but it was accurate information for audio-critical work.

The with-audio gap narrowed in version 1.1. See Happy Horse 1.1 for current benchmark data.


Version 1.0 Capabilities

The Transfusion architecture and general capability overview are in the Happy Horse family guide. The notes below focus on what version 1.0 achieved in each mode at launch.

Text-to-Video

Version 1.0 reached ~87% multi-shot narrative consistency at launch, the highest documented for any publicly available model at the time. This follows directly from the shared-parameter architecture: character, environment, and lighting information is processed in the same 32-layer shared space throughout generation, so context persists across shots rather than being re-established at each cut.

Image-to-Video

Happy Horse 1.0's strongest benchmark position at launch was I2V: an Elo of 1,391–1,414 in the image-to-video no-audio category set a new all-time record on Artificial Analysis. The source image is treated as part of the same unified token sequence processed alongside the text prompt — not as a separate conditioning signal — which is why visual identity holds continuously through generation rather than being approximated at a conditioning step.

Reference-to-Video

Accepts up to nine reference images alongside the text prompt for character or style anchoring, plus audio references for voice tone. The primary mode for production workflows requiring consistent character identity across multiple generated clips.

Video Editing

The video-edit endpoint accepts an existing video clip alongside a natural language description of desired changes and generates an edited version implementing those modifications. This enables post-generation refinement: changing backgrounds, adjusting lighting, or modifying character action, without regenerating from scratch.


Lip Sync: 14.60% WER

Happy Horse 1.0 supports native phoneme-level lip sync across seven languages (English, Mandarin, Cantonese, Japanese, Korean, German, French) with a documented WER (Word Error Rate) of 14.60%, the lowest of any publicly available model in 2026.

WER measures the percentage of words where the generated mouth movement does not accurately match the audio content. At 14.60%, the lip sync is achieved without any post-processing synchronization step: audio and visual tokens are generated jointly in the same forward pass, so the sync is inherent in the output rather than applied afterward.

Cantonese is listed as a distinct language separate from Mandarin, reflecting the Taotian Group's e-commerce base across Hong Kong, Guangdong, and diaspora communities.


How Happy Horse 1.0 Compared at Launch

vs. Seedance 2.0

Happy Horse 1.0 led on visual generation quality (no-audio Elo), generation speed, and multi-shot narrative consistency (~87%). Seedance 2.0 led on audio quality in blind tests and multimodal reference richness. For volume production and narrative work where speed and visual consistency were primary, Happy Horse 1.0 had the edge. For audio-critical content, Seedance 2.0 was the stronger choice.

vs. Veo 3.1

Veo 3.1 led on audio-visual synchronization precision, 4K resolution, and Scene Extension for longer-format content. Happy Horse 1.0 led on generation speed, visual quality in blind comparison, multi-shot consistency, and open licensing flexibility.

vs. Kling 3.0

Kling 3.0 offered longer single-pass clip duration and native 4K. Happy Horse 1.0 led on visual benchmark scores and generation speed. For production work at 1080p where quality and speed mattered, Happy Horse 1.0 was the stronger benchmark performer.


Known Limitations in 1.0

Motion at complex actions

At high-speed or physically demanding actions (running, fighting, rapid camera tracking), 1.0 occasionally produced motion that felt sluggish or floating rather than fully grounded. This was addressed as the primary motion improvement in version 1.1.

Subject drift during generation

Faces, logos, and objects could gradually shift in appearance during a clip, particularly in longer sequences. Temporal consistency improvements in 1.1 significantly reduced this behavior.

15-second maximum clip duration

Longer content requires chaining individual generations. No native Scene Extension or Video Extend feature was included in 1.0.

1080p maximum resolution

4K output is not available. For broadcast-grade 4K, Kling 3.0 and Veo 3.1 are the current alternatives.

With-audio benchmark gap

In the Artificial Analysis with-audio T2V category, Happy Horse 1.0 sat approximately 14 points behind Seedance 2.0. For most production use cases this gap was not decisive, but it was accurate context for audio-critical work.


Frequently Asked Questions

Should I still use Happy Horse 1.0?

For most new work, Happy Horse 1.1 is the better choice. It improves on 1.0's specific weaknesses (motion expressiveness and temporal consistency) with no change in generation speed or output specifications. Happy Horse 1.0 remains available if you have an existing pipeline calibrated to its specific output characteristics.

Does Happy Horse 1.0 support video editing?

Yes. The video-edit endpoint accepts an existing video clip alongside a natural language description of desired changes and generates an edited version. This is distinct from reference-to-video: the source material is a video, not a set of reference images.

What made Happy Horse 1.0's I2V performance notable?

The image-to-video Elo of 1,391–1,414 at launch set an all-time record at that point on Artificial Analysis. The source image is treated as part of the same unified token sequence the model processes alongside the text prompt, rather than as a conditioning signal applied to a separate generation process. The visual identity is maintained because the model reasons about it continuously throughout generation.

Why did Happy Horse 1.0 appear anonymously on the leaderboard?

No official explanation was provided by Alibaba. The anonymous entry followed by a three-day identification period and a deliberate public reveal generated significantly more coverage than a standard product launch would have — an outcome that may or may not have been the intended effect.


For the current production version: Happy Horse 1.1 · For the family overview: Happy Horse