Open Weights

MiniMax H3

One model for every modality. Combine text, images, video, and audio in a single context to generate 2K video with native stereo audio, up to 15 seconds.

  • 4.9/5 from 32K+ creators
  • 1.2M+ clips generated
  • Free — no account needed

Free and anonymous. Protected by Cloudflare Turnstile — no account, no card.

Made with MiniMax H3

One Context. Every Modality. Precise Control Throughout.

Reference to Video

Text, Images, Video, and Audio in One Context

Pass up to 9 images, 3 video clips, and 3 audio tracks together in a single generation. It reads identity, performance, camera movement, composition, soundscape, and editing rhythm from whatever you give it, then carries all of it through to one coherent result.

Precise Editing

Change One Thing, Keep Everything Else

Swap a product, rewrite signage, replace a line of dialogue, relight a scene from day to night, or add and remove objects. Localized edits land where you asked for them while the rest of the shot stays stable, so you can keep iterating on footage you already like.

Text and Interfaces

Typography, UI, and Graphics That Hold Together

H3 renders legible type, credits, subtitles, and brand marks, and it animates real interface work: product landing pages, game menus and HUDs, motion posters, and dynamic typography. Prompts run up to 7,000 characters, so a full shot list stays in one request.

Native Audio

Sound Composed to Picture

Every generation returns native stereo audio: original score, dialogue, foley, and room tone timed to the cut. Give MiniMax H3 a reference recording and it will transfer or clone that voice onto your character while keeping the performance intact.

How it works

Describe the shot in plain language, attach whatever reference material you already have, and check the opening frame before committing to the full render. Reviewing a still first is the cheapest way to catch a framing or lighting problem, because fixing it costs one more sentence rather than a second clip.

  1. 1

    Describe the shot

    Enter up to 7,000 characters describing the subject, camera, lighting, soundscape, and editing rhythm.

  2. 2

    Add references

    Pass up to 9 images, 3 video clips, and 3 audio tracks to set the style, identity, and motion exactly how you want it.

  3. 3

    Generate opening frame

    Review the generated opening frame to ensure the composition and lighting match your vision perfectly.

  4. 4

    Animate to 2K with stereo audio

    Confirm the generation and watch the scene come to life in 2K resolution with synchronized native audio.

Examples

What creators are building with MiniMax H3

Brand films, character work, stylized animation, gameplay, and video editing, all from the same model.

First and Last Frame

Vintage Binocular Brand Film

Use Images 1-4 as sequential keyframes, seen through a vintage binocular viewfinder searching for the MINIMAX installation. Open out of focus with subtle handheld shake, then push in quickly and rack focus onto Image 1. Between keyframes, use fast binocular-scan transitions with whip movement, motion blur, optical smearing, and brief exposure flicker. Keep the twin circular lens mask absolutely fixed throughout. Visual language: a voyeuristic, Wes Anderson-inspired 35mm film look with fine grain, soft highlight halation, restrained color, and red typographic accents.
Reference to Video

Fantasy Wuxia Character Film

Use Image 2 as the locked character reference. Preserve the half-up long black hair, openwork silver crown, indigo ribbon, layered pale hanfu, translucent blue outer robe, deep-blue sash, silver floral fastener, and long tassels. Use Image 1 for storyboard order and pacing. Render in high-quality 4K, 16:9 Chinese-inspired 3D with cinematic xianxia production value: intense, solemn, and shaped by destiny. Follow the storyboard beat by beat, with natural camera movement and seamless transitions.
Video Reference

Live-Action to Voxel Transformation

Preserve the buildings, pedestrians, and overall environment in Video 1 as photoreal live action. Transform only the trees and cars into 3D pixel-art or voxel-block objects in the style of Minecraft, using Image 1 as the visual reference. Keep their motion physically correct, and preserve the real environment's shadows and transmitted light. Use Video 2 as the overall target.
Text to Video

Neon Laundromat Encounter

15 seconds, 16:9 landscape. Combine a live-action late-night laundromat with hand-drawn luminous animation. The small self-service laundromat has gently flickering fluorescent lights, running washers, plastic baskets, a worn bench, and one sock on the floor. Keep the space quiet and faintly nostalgic. Use a one-handed phone-camera feel with visible shake, exposure fluctuation under white fluorescent light, environmental reflections in glass, and delayed autofocus at close range. Avoid polished commercial composition; it should feel like an authentic late-night encounter, filmed while following a strange apparition.
Motion Reference

Capybara Motion Recreation

Match the action in Video 1 from a locked-off wide camera. Replace the three suited men with three highly photoreal capybaras. Preserve the original movement path exactly: all three drop quickly to the floor; the left capybara jumps to center; the center capybara rolls to the far left; the new center capybara rolls to the far right; the right capybara jumps to center; finally, the center capybara jumps onto the other two, forming a pyramid. Keep the camera fixed and integrate fur, lighting, and shadows realistically into the scene.
Video Editing

Green-Screen Fairytale Composite

Remove the green screen background of Video 1 and turn it into a fairy tale-like background similar to Video 2. The background elements need to completely match the actions of the characters in Video 1. Modify the lighting of the characters in Video 1 so that it completely matches the background.

The numbers behind every generation

Output is 2K by default, which puts 1440 pixels on the short edge for ratios between 16:9 and 9:16 and reaches roughly 3.7 megapixels on wider formats — 2976x1248 at 21:9. Clips run 5 to 15 seconds at 24 frames per second, and every one of them comes back with a synchronized stereo mix rather than a silent picture you have to score afterwards.

Resolution (2976x1248)
2KResolution (2976x1248)
Duration at 24 FPS
5–15sDuration at 24 FPS
Stereo Audio
NativeStereo Audio
Max References
12Max References
Aspect Ratios
7Aspect Ratios
Image Refs
9Image Refs
Character Prompts
7,000Character Prompts
Open Weights
YesOpen Weights

How MiniMax H3 compares

MiniMax H3 compared with Sora 2, Veo 3.1, Kling 3 and Hailuo 02 across resolution, duration, native audio, reference budget and open weights.
FeatureMiniMax H3Sora 2Veo 3.1Kling 3Hailuo 02
Resolution2K (2976x1248)1080p1080p1080p720p
Duration5–15sUp to 60s5s5–10s5s
Native AudioYes (Stereo)NoNoNoNo
Reference Budget12 files (9 img, 3 vid, 3 aud)1 image1 image1 image1 image
Open WeightsYesNoNoNoNo

The gap that matters is not raw resolution — several of these models reach a comparable pixel count. It is that the rivals treat text-to-video, image-to-video, motion reference and audio as separate endpoints, so a shot that needs three of them needs three round trips and a compositor to reconcile the results. Passing a single unified context, and getting a scored stereo mix back with the picture, is what removes that step.

Trusted by leading creators

MiniMax H3 completely changed our post pipeline. Having one context that understands video references alongside audio saves us days of manual composite work.
Sarah JenkinsVFX Supervisor, Studio North
The native stereo audio generation is indistinguishable from magic. It does not just make sound, it understands the physical space of the shot we prompt.
David ChenSound Designer, Waveform
Finally, a model that renders typography cleanly. We use the UI/UX motion capability to prototype entire app flows directly from Figma frames.
Elena RostovaCreative Director, Form and Function

Common questions about MiniMax H3

MiniMax H3 is an open-weights, general-purpose multimodal video model. Instead of a separate model for each task, MiniMax H3 reads text, images, video, and audio in one unified context and generates coherent audiovisual results from any mix of them. It supports text-to-video, first-and-last-frame, reference-to-video, and precise video editing.

Ready to generate 2K video?

Experience the unified context, native audio, and open-weights freedom of MiniMax H3 today.

Try Free