MiniMax H3: 2K Video, Native Audio & Multimodal Input
MiniMax H3 guide: 2K video, native stereo audio, multimodal inputs, first and last frame control, and the VioEvo workflow.
Tags
Create a MiniMax H3 video on VioEvo.
Developer: MiniMax · Released: July 31, 2026 · Category: Omni-modal video generation · Official release: MiniMax H3 · Available on VioEvo: Text to Video and Image to Video
What Is MiniMax H3?
MiniMax H3 is a video generation system built around a single goal: interpret richer creative context before generating a short audiovisual clip. MiniMax describes H3 as an omni-modal model that can jointly understand text, images, video, and audio. Its output combines motion with native stereo audio, rather than treating sound as a separate post-production step.
That positioning matters most when a request has more than one constraint. A product shot may need a particular final composition, a specific camera move, a visible action, and sound that belongs to the scene. H3 is designed to reason across those relationships before it produces the video.
For VioEvo creators, the practical entry points are simpler: write a text prompt for a new clip, or provide a start image with an optional end image to guide the motion and final composition.
MiniMax H3 Specs at a Glance
| Capability | Current specification |
|---|---|
| Release date | July 31, 2026 |
| Clip duration | 4 to 15 seconds |
| Output options | 768P and 2K |
| Output frame rate | 24 FPS |
| Audio | 32 kHz stereo |
| Official input modes | Text to video, first-frame or last-frame image to video, first-and-last-frame image to video, and reference to video |
| VioEvo modes | Text to Video; Image to Video with one start frame or an optional end frame |
MiniMax documents 768P as the base output and describes a separate H3 regeneration stage for 2K results. Its public video API exposes both 768P and 2K output choices, with a required duration between 4 and 15 seconds.
Sources: MiniMax's H3 launch announcement and the MiniMax video-generation API reference.
The H3 Workflow: Context, Generation, and 2K
H3 is best understood as a workflow with three distinct jobs, rather than a single resolution toggle.
Context understanding. MiniMax's H3-Context-IR service can inspect text, images, video, and audio together, then turn the request into a structured, semantically enriched video prompt. It is intended to preserve the creative request while resolving how the provided inputs relate to one another.
768P audiovisual generation. The H3 base generation stage produces the initial clip with sound and picture together. This is the practical iteration point: evaluate framing, motion, character behavior, and audio-visual fit before moving to a higher-resolution deliverable.
2K regeneration. MiniMax describes H3's high-resolution path as regeneration of the lower-resolution result in context, rather than a conventional standalone super-resolution step. This is why 2K is best treated as a final-output decision after the creative direction is working.
This does not guarantee that every complex request will be resolved perfectly. It does give H3 a clear use case: short clips where visual direction and scene audio need to agree.
Source: MiniMax H3 release announcement.
Using MiniMax H3 on VioEvo
VioEvo currently exposes the two H3 workflows that are most useful for everyday creation:
- Text to Video: start from a written shot description and choose a duration from 4 to 15 seconds.
- Image to Video: supply one start image, then optionally add an end image when the intended final composition matters.
- Resolution: choose 768P while refining a concept, then select 2K when the composition and motion are ready for a higher-resolution delivery.
- Audio: H3 generation on VioEvo includes generated audio, so describe the desired dialogue, ambience, or sound effects in the prompt instead of assuming the clip should be silent.
The official MiniMax API also includes reference-to-video input, but this guide covers the Text to Video and Image to Video H3 paths currently available in VioEvo. That distinction is important: a feature described in a model API is not automatically a feature exposed in every creative tool.
When to Use Text to Video
Choose Text to Video when the scene can be expressed as a single directed shot: a product reveal, a cinematic establishing shot, a short action beat, or a speaker in a defined setting. H3 is a strong candidate when audio direction is part of the idea, such as a line of dialogue, room tone, a specific environmental sound, or a music cue that should match the action.
Start with a short 768P test clip. It is usually easier to correct camera direction, action timing, or sound design before committing the idea to a higher-resolution output.
When to Use Image to Video
Choose Image to Video when the opening composition is non-negotiable: a product photograph, approved key art, a character design, or a carefully styled scene. Add an end frame when the finishing layout matters too, for example, when a product must land centered in a specific pose or a transition must end on a defined title composition.
The image frames guide the boundaries of the shot. The text prompt should still explain what happens between them: camera movement, subject action, lighting behavior, and sound.
How to Prompt MiniMax H3
H3's official API requires a text instruction even when images or other reference media are present. Treat that text as the shot brief, not a bag of keywords.
Use this structure:
[Shot type] of [subject] in [setting]. The camera [movement] while [action].
[Lighting and visual treatment]. Sound: [dialogue, ambience, or effects].
[Duration] seconds, [aspect ratio].
For example:
Medium tracking shot of a red electric scooter moving through a rain-slick city street at night.
The camera keeps pace from the side as reflections stretch beneath the wheels.
Neon storefronts and soft mist, realistic commercial lighting. Sound: electric motor hum,
light rain, and distant traffic. 8 seconds, 16:9.
Keep the description internally consistent. A prompt asking for a locked-off close-up, a sweeping aerial move, two different locations, and three competing sound sources in a few seconds creates an ambiguous shot plan. A focused motion beat and one clear audio environment give the model a more coherent brief.
Is MiniMax H3 Fully Open Source?
No. MiniMax positions H3 as an open model, but the complete high-resolution workflow is not fully open-source. The company's documentation explicitly states that the implementation of H3-Context-IR is not open sourced. That service returns an enhanced prompt; it does not create a video by itself.
The practical takeaway is straightforward: distinguish between access to H3 model materials and the availability of every component in MiniMax's production pipeline. For creators using VioEvo, this is mostly an implementation detail. For teams evaluating self-hosting or reproducing MiniMax's full 2K workflow, it is a material constraint.
Source: MiniMax H3-Context-IR API documentation.
MiniMax H3 FAQ
How long can a MiniMax H3 video be?
MiniMax H3 supports durations from 4 to 15 seconds. VioEvo exposes that same range for its H3 Text to Video and Image to Video workflows.
Does MiniMax H3 generate audio?
Yes. MiniMax specifies native 32 kHz stereo audio, and VioEvo's H3 generation includes audio. Describe the intended sound in the prompt when it matters to the shot.
Can I control the first and last frames?
Yes. On VioEvo's Image to Video workflow, H3 accepts a required start image and an optional end image. Use two images when you need to guide both the opening and closing composition.
Is 2K just an upscale?
MiniMax describes its 2K path as in-context regeneration of an H3 768P result, not a conventional standalone super-resolution module. The useful production decision remains the same: validate the shot at 768P, then choose 2K for a finished output when higher resolution is needed.
Create a MiniMax H3 Video
Start with a concise shot brief in Text to Video. When the beginning or ending composition is important, switch to Image to Video and use one or two frame images to anchor the result.